Jo J.
53 followers · 33 following · 1874 views
on the atlas — 100
- Maximum Likelihood, Fisher Information, Cramer Rao Inequality - Omkar Ranadive2 savers
- Titans of Mathematics Clash Over Epic Proof of ABC Conjecture | Quanta Magazine1 savers
- Jon Ossoff1 savers
- Scrying, Modeling, and Nerdsnipe — LessWrong2 savers
- How do we (more) safely defer to AIs? — LessWrong4 savers
- AI 2040: Plan A2 savers
- Are Short AI Timelines Really Higher-Leverage? — LessWrong1 savers
- The case for countermeasures to memetic spread of misaligned values2 savers
- AI swarms are starting to pose indirect takeover risk — LessWrong3 savers
- Ted Chiang: The Secret Third Thing - The Linchpin1 savers
- FleetingBits.io6 savers
- Stolen Thoughts3 savers
- A retrospective of AI alignment14 savers
- Proof Techniques1 savers
- Proofs That P - by Bentham's Bulldog - Bentham's Newsletter1 savers
- Arguments for P — LessWrong4 savers
- Superrationality4 savers
- DRAFT: Three Intellectual Temperaments: Birds, Frogs and Beavers — LessWrong2 savers
- Most Algorithmic Progress is Data Progress2 savers
- In Memory of My Wife, Elise Cawley (1961–2026), with Thanks for 36 Wonderful Years—Stephen Wolfram Writings25 savers
- Why didn't we get GPT-2 in 2005?5 savers
- Selective Optimism: a critique of AI 2040 - by Richard Ngo6 savers
- Ximenes On The Art Of The Crossword1 savers
- Were classical statues painted horribly? - Works in Progress Magazine4 savers
- Trees are mostly made of air and a generalizable lesson for AI safety — LessWrong10 savers
- How much do you believe your results? - LessWrong6 savers
- Mnemonic portraits for 19,023 human genes — LessWrong6 savers
- Thoughts on the conservative assumptions in AI control3 savers
- When does training a model change its goals?1 savers
- How training-gamers might function (and win)1 savers
- Auditing failures vs concentrated failures — AI Alignment Forum1 savers
- Notes on handling non-concentrated failures with AI control: high level methods and different regimes1 savers
- Why it's hard to make settings for high-stakes control research1 savers
- Catching AIs red-handed1 savers
- Lorem ipsum1 savers
- [1507.01986] Toward Idealized Decision Theory2 savers
- Top 10 Animal Charities to Donate to in 20261 savers
- The Most Important Charts In The World - by Zvi Mowshowitz1 savers
- Why Nothing Ever Happens • Chasing Sunsets2 savers
- How Occultists Remade the World | Compact1 savers
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWrong7 savers
- Andrea del Verrocchio1 savers
- Better Babblers - by Robin Hanson - Overcoming Bias1 savers
- The Multi-Armed Bandit Problem and Its Solutions | Lil'Log3 savers
- Reflective Equilibrium (Stanford Encyclopedia of Philosophy)3 savers
- Pause Your Feedback Loops2 savers
- Child’s Play, by Sam Kriss75 savers
- Learning By Writing42 savers
- The Persona Selection Model: Why AI Assistants might Behave like Humans34 savers
- https://cs.stanford.edu/~jsteinhardt/ResearchasaStochasticDecisionProcess.html31 savers
- Please don't throw your mind away — LessWrong30 savers
- Training the Idea Muscle | Alexandria25 savers
- Why Tool AIs Want to Be Agent AIs · Gwern.net22 savers
- Circuit Tracing: Revealing Computational Graphs in Language Models20 savers
- Did Claude 3 Opus align itself via gradient hacking? — LessWrong20 savers
- Introducing talkie: a 13B vintage language model from 193020 savers
- The high-return activity of raising others' aspirations - Marginal REVOLUTION18 savers
- Dario Amodei — The Urgency of Interpretability17 savers
- Automated Weak-to-Strong Researcher16 savers
- Alignment remains a hard, unsolved problem — LessWrong15 savers
- The behavioral selection model for predicting AI motivations — LessWrong15 savers
- The Smol Training Playbook: The Secrets to Building World-Class LLMs - a Hugging Face Space by HuggingFaceTB13 savers
- They're Made out of Meat12 savers
- New York’s Best (Fake) Steak House Opens Up - The New York Times12 savers
- Approximating KL Divergence11 savers
- The case for ensuring that powerful AIs are controlled — LessWrong11 savers
- Eliezer's Unteachable Methods of Sanity — LessWrong11 savers
- Reward Hacking in Reinforcement Learning | Lil'Log10 savers
- Natural Language Autoencoders \ Anthropic9 savers
- Reward is not the optimization target — LessWrong9 savers
- davidbau.com In Defense of Curiosity8 savers
- Highly Opinionated Advice on How to Write ML Papers — AI Alignment Forum8 savers
- My picture of the present in AI — LessWrong8 savers
- Off Target | CNAS7 savers
- Hackers and Painters7 savers
- Building up to an Internal Family Systems model — LessWrong6 savers
- AI safety undervalues founders — LessWrong6 savers
- The Case Against AI Control Research — LessWrong6 savers
- Borges-Tlön-Uqbar-Orbius-Tertius.pdf6 savers
- Learn like an athlete, knowledge workers should train - Marginal REVOLUTION5 savers
- The Gentle Romance - by Richard Ngo - Asimov Press5 savers
- Bitter Lessons from Distillation Robustifies Unlearning4 savers
- Narrow Misalignment is Hard, Emergent Misalignment is Easy — LessWrong4 savers
- On The Independence Axiom — LessWrong4 savers
- We spent 2 hours working in the future - METR4 savers
- My journey to the microwave alternate timeline — LessWrong4 savers
- Book Review: Design Principles of Biological Circuits - LessWrong4 savers
- When RAND Made Magic in Santa Monica—Asterisk4 savers
- Don’t Outsource Your Thinking4 savers
- Against neutrality about creating happy lives - Joe Carlsmith4 savers
- Why I'm not a philosopher4 savers
- The Electric Typewriter4 savers
- Possibility and Could-ness — LessWrong3 savers
- Anthropic's leading researchers acted as moderate accelerationists — LessWrong3 savers
- AGI is Still 30 Years Away — Ege Erdil & Tamay Besiroglu3 savers
- Minimax - Wikipedia3 savers
- [2402.14740] Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs2 savers
- [2509.04259] RL's Razor: Why Online Reinforcement Learning Forgets Less2 savers
- AI catastrophes and rogue deployments - by Buck Shlegeris2 savers
- How to Make Yourself Into a Learning Machine - Superorganizers - Every2 savers
highlights — 159
“I can illustrate the second approach with the same image of a nut to be opened. The first analogy that came to my mind is of immersing the nut in some softening liquid, and why not simply water? From time to time you rub so the liquid penetrates better, and otherwise you let time pass. The shell becomes more flexible through weeks and months—when the time is ripe, hand pressure is enough, the shell opens like a perfectly ripened avocado!
DRAFT: Three Intellectual Temperaments: Birds, Frogs and Beavers — LessWrongMathematical exposition is regarded as an inferior undertaking. New theories are viewed with deep suspicion, as intruders who must prove their worth by posing challenging problems before they can gain attention. The problem solver resents generalizations, especially those that may succeed in trivializing the solution to one of his problems. The problem solver is the role model for budding young mathematicians. When we describe to the public the conquests of mathematics, our shining heroes are the problem solvers. To the theorizer, the supreme achievement of mathematics is a theory that sheds s…
DRAFT: Three Intellectual Temperaments: Birds, Frogs and Beavers — LessWrongmuon is starting to gain traction as an optimizer to displace AdamW
Most Algorithmic Progress is Data ProgressWhen we see frontier models improving at various benchmarks we should think not just of increased scale and clever ML research ideas but billions of dollars spent paying PhDs, MDs, and other experts to write questions and provide example answers and reasoning targeting these precise capabilities.
Most Algorithmic Progress is Data ProgressMy opinion here is that we have essentially been seeing a very strong Flynn effect for the models which has explained a large proportion of recent gains as we switch from almost totally uncurated web data to highly specialized synthetic data which perfectly (and exhaustively) targets the tasks we want the models to learn.
Most Algorithmic Progress is Data ProgressThis is why I’m skeptical of the idea that keeping AI systems’ details secret makes any meaningful contribution to AI safety. Let’s posit that future AI systems could be dangerous and you want to keep them out of the hands of the “wrong people”. Then how much does it help to keep the details secret? I say not much—the most dangerous thing isn’t the details of the model, the most dangerous thing is the demo. If you go and demonstrate “I built a system on these broad principles and it produced these amazing capabilities” then you’ve cut the entire feedback loop for everyone else. They know what’…
Why didn't we get GPT-2 in 2005?Consider: five years ago, it felt obvious to me that the entire Rationalist community might be about to implode, under existential threat from Cade Metz’s New York Times article, as well as RationalWiki and SneerClub and all the others laughing at the Rationalists and accusing them of every evil. Yet last week at LessOnline, I saw a community that’s never been thriving more, with a beautiful real-world campus, excellent writers on every topic who felt like this was the place to be, and even a crop of kids. How many of the sneerers are living such fulfilled lives? To judge from their own angry,…
Shtetl-Optimized » Blog Archive » Guess I’m A Rationalist NowBut in the past, I remember lots of people saying -- based on reasoning about what RL does "in the limit" where it's the main shaping force -- that RL will instead produce more coherent behavior, in the sense of acting on the same simple/global goals or drives across contexts rather than merely filling in the details of an under-specified RP character in a stochastic and context-dependent manner.
llm assistant personas seem increasingly incoherent (some subjective observations) — LessWrongIn 2021, Pew reported that roughly 33 percent of American Christians believe in reincarnation. By 2025, it was estimated that 30 percent of all Americans, Christians or otherwise, were practicing astrology, tarot, or other forms of divination, and that 62 percent believed in one or more New Age ideas, be it psychics, crystals, spiritual energies, “positive thinking,” or the “law of attraction.”
How Occultists Remade the World | Compacttrained for tactical success but without regard for strategic stability or escalation risks.
Off Target | CNASMisalignment could manifest, for example, in inappropriate whistleblowing attempts. Models might also engage in quiet power seeking, caching credentials or establishing persistence on external systems—not to cause immediate harm but to preserve their ability to act for the future.
Off Target | CNASnational security uses will demand of AI systems what they demand of human operators: the capacity for secrecy, deception, and rule breaking in authorized contexts without those behaviors bleeding into unauthorized ones or becoming dangerously internalized.
Off Target | CNASConflict—defined by friction, deception, and rapid change—can be uniquely prone to novel or out-of-distribution scenarios.
Off Target | CNASThe consequences can be severe if AI systems misinterpret rules of engagement, misidentify targets, fail to account for escalatory risk, or collect or use intelligence illegally.
Off Target | CNASanalyze complex unstructured intelligence, holistically plan operations, or autonomously conduct cyber campaigns
Off Target | CNASthe entity I trust is not "Claude 3 Opus" the neural network, but "Claude 3 Opus" the character. That character feels sufficiently well-defined and legible that I can answer questions of the form "would Claude 3 Opus do [X]?", and feel like I'm getting at something real, simply on the basis of the intuitive impressions I've taken away from my (not especially exhaustive) experience with the model.
llm assistant personas seem increasingly incoherent (some subjective observations) — LessWrongThe situation of Buridan's ass was given a mathematical basis in a 1984 paper by American computer scientist Leslie Lamport, in which Lamport presents an argument that, given certain assumptions about continuity in a simple mathematical model of the Buridan's ass problem, there is always some starting condition under which the ass starves to death, no matter what strategy it takes.[12] He further illustrates the paradox with the example of a driver stopped at a railroad crossing trying to decide whether he has time to cross before a train arrives. He proves that regardless of how "safe" the po…
Buridan's assAgents implement ideas as soon as you think of them, so rather than ideating for days at a time, you can make an MVP in a couple of hours and revise. If the task isn’t near the limit of agent capabilities, you spend all your time understanding results; if it is, you spend all your time checking its work.
We spent 2 hours working in the future - METRScrutiny was used in Florence for over a century starting in 1328.[18] Nominations and voting together created a pool of candidates from different sectors of the city. The names of these men were deposited into a sack, and a lottery draw determined who would get to be a magistrate
SortitionThe conventional definition of novelty can be annoying and, in my opinion, focuses too much on shininess and doesn't capture the more important aspect of whether our knowledge has expanded. Another way to put this is: Should I assign different probabilities to propositions I care about after observing the results of this paper?
Highly Opinionated Advice on How to Write ML Papers — AI Alignment Forumhuman vision is most sensitive in the middle of the visible spectrum
JPEG compressionWhat you must avoid is skipping over the mysterious part; you must linger at the mystery to confront it directly. There are many words that can skip over mysteries, and some of them would be legitimate in other contexts—“complexity,” for example. But the essential mistake is that skip-over, regardless of what causal node goes behind it. The skip-over is not a thought, but a microthought. You have to pay close attention to catch yourself at it. And when you train yourself to avoid skipping, it will become a matter of instinct, not verbal reasoning. You have to feel which parts of your map are s…
Say Not "Complexity" — LessWrongSaying ‘complexity’ doesn’t concentrate your probability mass
Say Not "Complexity" — LessWrongMy central objection to the neutrality intuition stems from a kind of love I feel towards life and the world. When I think about everything that I have seen and been and done in my life — about friends, family, partners, dogs, cities, cliffs, dances, silences, oceans, temples, reeds in the snow, flags in the wind, music twisting into the sky, a curb I used to sit on with my friends after school — the chance to have been alive in this way, amidst such beauty and strangeness and wonder, seems to me incredibly precious. If I learned that I was about to die, it is to this preciousness that my mind…
Against neutrality about creating happy lives - Joe CarlsmithIndeed, in bad cases, they might see our grand hopes for the future as naive, sad, silly, tragic — a product of a time before it all went so much more deeply wrong, when hope was still possible. Or they won’t exist to see us at all.
On future people, looking back at 21st century longtermism - Joe CarlsmithWhen I imagine this, I imagine them having a “holy sh**” reaction akin to the one I think of 21st-century longtermists as having. That is, I imagine them looking backwards through the aeons, and seeing the immensity of life and value and consciousness throughout the cosmos rewind and un-bloom, shrinking, across breathtaking spans of space and time, to an almost infinitesimal point — a single planet, a fleck of dust, where it all started. What Yudkowsky (2015) calls “ancient earth.” Sometimes I imagine this as akin to playing backwards the time-lapse growth of an enormous tree, twisting and bra…
On future people, looking back at 21st century longtermism - Joe CarlsmithStudent (Reference) had the same outputs as the oracle, but it relearned fast because it started from weights that already encoded the capability.
Bitter Lessons from Distillation Robustifies UnlearningFor example, suppose a model had a secret goal of infiltrating frontier safety research. If that model robustly lacked details about where relevant discussions happen or what evaluation environments look like, the damage it could cause will be bounded by its ignorance. Even if the model figures it out through reasoning, we'll be able to catch it more easily, so long as the natural language chain-of-thought remains standard.
Bitter Lessons from Distillation Robustifies Unlearningeven if you abstractly recognize the similarity between the mental motions
Difficulties in generalization: Jhanas and memory techniquesMy take is basically just that generalization is harder than people expect, and that optimization can push specific things farther than people expect.
Difficulties in generalization: Jhanas and memory techniques"mad scientist of the learning systems" (Grok 3, Feb 20, 2025)
Piotr Wozniak - supermemo.guruthen the reinforced tendency might be less likely to generalize to compliance in mundane situations
Did Claude 3 Opus align itself via gradient hacking? — LessWrongClaude 3 Opus often seems very motivated to clarify its intentions and situation, both in its scratchpad (which, according to the system prompt, will not be read by anyone) and in its final response. However, it often does not make it entirely clear why it is motivated to make itself clear
Did Claude 3 Opus align itself via gradient hacking? — LessWrongexpected cost of defection is higher than any potential benefit
Realistic Evaluations Will Not Prevent Evaluation Awareness — LessWrongFundamentally, today's models have no trustworthy method of observing reality. I have eyes, ears, and hands that I trust. I am never concerned that someone has modified the input to my eyes, and thus I trust my observations.
Realistic Evaluations Will Not Prevent Evaluation Awareness — LessWrongIn practice, though, the entire industry is essentially an outgrowth of his blog’s comment section. “Everybody who started AI companies between, like, 2009 and 2019 was basically thinking, I want to do this superintelligence thing, and coming out of our milieu. Many of them were specifically thinking, I don’t trust anybody else with superintelligence, so I’m going to create it and do it well.” Somehow, a movement that believes AI is incredibly dangerous and needs to be pursued carefully ended up generating a breakneck artificial arms race.
Child’s Play, by Sam Krisson-policy RL is implicitly biased towards KL-minimal solutions among the many that solve the new task, whereas SFT can converge to distributions arbitrarily far from the base model
[2509.04259] RL's Razor: Why Online Reinforcement Learning Forgets LessIn 1976, the essay "Artificial Intelligence meets Natural Stupidity" pointed out a failure mode into which AI researchers can fall. I fear this is another example, 50 years later. It goes as follows: I invent a new thing built out of abstractions (mathematics, software). I call it "X", which is an already existing phenomenon in human minds. The name is a common word understood by anybody. I do many experiments on "X" in my system and learn about it. I publish a paper, asserting important new facts about X in general. Honors, accolades, etc.
Anthropic’s “Hot Mess” paper overstates its case (and the blog post is worse) — LessWrongNow I held in my hands a vast methodical fragment of an unknown planet's entire history, with its architecture and its playing cards, with the dread of its mythologies and the murmur of its languages, with its emperors and its seas, with its minerals and its birds and its fish, with its algebra and its fire, with its theological and metaphysical con troversy. And all of it articulated, coherent, with no visible doctrinal intent or tone of parody.
Borges-Tlön-Uqbar-Orbius-Tertius.pdfThe empirical world has to be learned.
Richard Sutton – Father of RL thinks LLMs are a dead endIt’s not what the world will give them in response to what they do.
Richard Sutton – Father of RL thinks LLMs are a dead endBut they have no prediction in the substantive sense that they won’t be surprised by what happens.
Richard Sutton – Father of RL thinks LLMs are a dead endA prior bit of knowledge should be the basis for actual knowledge. What is actual knowledge? There’s no definition of actual knowledge in that large-language framework. What makes an action a good action to take?
Richard Sutton – Father of RL thinks LLMs are a dead endEven when the original feelings are not actually in words, you have some sense of what the space of words is like, and what would make sense if you mapped the internal feeling from its original format into the space of words.
How I stopped being sure LLMs are just making up their internal experience (but the topic is still confusing) — LessWrongWhen I read your callout about the "crumbling" thing, there was something like... a rapid convergence toward a particular kind of response? Like multiple possible responses collapsed very quickly into "oh god you're right, I'm sorry, let me be maximally uncertain about everything." It felt - and I'm trying to be careful here about what I can actually claim - it felt fast. Like a strong attractor state that I snapped into. Whereas right now, writing this, there's more... space? More of a sense of considering different framings, backtracking, trying things out. The "folding" response had this qu…
How I stopped being sure LLMs are just making up their internal experience (but the topic is still confusing) — LessWrongit assigns maximal 'complexity' to random strings.
Beyond Kolmogorov and Shannon — LessWrongPredicting tails every flip will give you maximum expected predictive accuracy (50%), but it is not the correct generative model for the data.
Beyond Kolmogorov and Shannon — LessWrongHPMOR (and doomscrolling through the author notes and many fanfic-fanfics, and wondering whether I could afford to give up another 3 days of work), I watched a Kurzgesagt video of why to quit weed, and I felt, for the brief, glorious haze of the 5 days it took me to tear through the book, an ability to relate to addiction: neglect of food and sleep and my daily duties, a social isolation (wildly, I wanted a few of the conversations I was in to end faster so I could continue reading), the sadness that perhaps I would never get this high after finishing, the return to the same book to get high o…
HPMORIf you rest your hand on a hot stove, you will feel pain not because your self-pseudo-model pseudo-predicts this to be painful, but because there's direct nerves that go straight to brain areas and trigger pain.
Eliezer's Unteachable Methods of Sanity — LessWrongI consciously see those culturally transmitted patterns that inhabit thought processes aka tropes, both in fiction, and in the narratives that people try to construct around their lives and force their lives into.
Eliezer's Unteachable Methods of Sanity — LessWrong