Celeste
5 followers · 1 following · 289 views
on the atlas — 27
- Writing Doom (2024) directed by Suzy Shepherd • Reviews, film + cast • Letterboxd1 savers
- Answer to Job | Slate Star Codex3 savers
- Self-Help Tactics That Are Working For Me — LessWrong2 savers
- Pain is not the unit of Effort38 savers
- Searching for outliers | benkuhn.net36 savers
- Ludic's Guide To Getting Software Engineering Jobs — Ludicity3 savers
- Discovering state-of-the-art reinforcement learning algorithms | Nature2 savers
- How to Twitter Successfully | near.blog11 savers
- trees are harlequins, words are harlequins — I don't think you're drawing the right lesson from...2 savers
- The “it” in AI models is the dataset. – Non_Interactive – Software & ML8 savers
- 2025 LLM Year in Review | karpathy16 savers
- Post 38: On Slack - Having room to be excited — Neel Nanda23 savers
- Half-assing it with everything you've got33 savers
- Ugh fields — LessWrong3 savers
- Make product worse, get money2 savers
- Identifying Functionally Important Features with End-to-End Sparse Dictionary Learning — LessWrong1 savers
- Slack Has Positive Externalities For Groups — LessWrong1 savers
- Strategies for learning - by Andy Masley1 savers
- Lies, Damn Lies, and Fabricated Options — LessWrong2 savers
- 'San Francisco is just a place' is not just a deepity1 savers
- The Plan (Now-June 2027) - by Lydia Nottingham - pronotre1 savers
- Area Man Angry AI Not Silver Bullet, by Gwern, Gemini-2.5-pro, GPT-4-o3 · Gwern.net1 savers
- https://cs.stanford.edu/~jsteinhardt/ResearchasaStochasticDecisionProcess.html31 savers
- Aman's AI Journal • Primers • Ilya Sutskever's Top 3020 savers
- Dario Amodei — The Urgency of Interpretability17 savers
- Preparing for the Intelligence Explosion | Forethought6 savers
- X explains Z% of the variance in Y — LessWrong5 savers
highlights — 5
I’ve always needed a “win condition”. I need to know it’s possible, even if only someday at the end of my life, to be confident and secure in the feeling that I’ve done a good job. I’m fond of triumph, of trumpet fanfares and celebrations, and I need to believe that under some conditions, if I did certain sufficiently awesome things, I’d have earned a triumphant celebration. A glorious state, free of worry, completely exultant. And, as a more immediately practical matter, I need conditions under which I’d let myself completely physically relax — constant straining literally hurts my back.
Self-Help Tactics That Are Working For Me — LessWrongBut who is actually going to build the LLM GUI? In this world view, nano banana is a first early hint of what that might look like. And importantly, one notable aspect of it is that it's not just about the image generation itself, it's about the joint capability coming from text generation, image generation and world knowledge, all tangled up in the model weights.
2025 LLM Year in Review | karpathyI've vibe coded entire ephemeral apps just to find a single bug because why not - code is suddenly free, ephemeral, malleable, discardable after single use. Vibe coding will terraform software and alter job descriptions.
2025 LLM Year in Review | karpathyRelated to all this is my general apathy and loss of trust in benchmarks in 2025. The core issue is that benchmarks are almost by construction verifiable environments and are therefore immediately susceptible to RLVR and weaker forms of it via synthetic data generation. In the typical benchmaxxing process, teams in LLM labs inevitably construct environments adjacent to little pockets of the embedding space occupied by benchmarks and grow jaggies to cover them. Training on the test set is a new art form.
2025 LLM Year in Review | karpathyThe subtlety with the Ugh Field is that the flinch occurs before you start to consciously think about how to deal with the Unhappy Thing, meaning that you never deal with it,
Ugh fields — LessWrong