Will Anderson
36 followers · 29 following · 1779 views
on the atlas — 281
- Conservation of Expected Evidence — LessWrong1 savers
- AI as orderly evacuation vs stampede - by Richard Ngo1 savers
- Doom as a bad method not a utopia trade-off — LessWrong3 savers
- How to have executive function – scribbles in the margins6 savers
- Proposal for tracking the effects of architecture on monitorability — Redwood Research2 savers
- Do not expect the unexpected | Ted Sanders3 savers
- The Commitment Races problem — AI Alignment Forum1 savers
- Deontology and virtue ethics as "effective theories" of consequentialist ethics — LessWrong4 savers
- Evolvable AI: Threats of a new major transition in evolution | PNAS1 savers
- The Rogue Agent Explosion Will Be Mostly Invisible — LessWrong1 savers
- Self Hosting — LessWrong1 savers
- Can a superintelligence do THAT? — LessWrong1 savers
- How Should the US Prepare for Increasingly Automated AI R&D? | IFP3 savers
- [2603.24676] When Is Collective Intelligence a Lottery? Multi-Agent Scaling Laws for Memetic Drift in LLMs1 savers
- What failure looks like - LessWrong12 savers
- New report: "Scheming AIs: Will AIs fake alignment during training in order to get power?" — AI Alignment Forum2 savers
- The behavioral selection model for predicting AI motivations — LessWrong15 savers
- Striving to Live a Perfectly Mundane Future - yeedrag1 savers
- Tools for keeping focused | benkuhn.net7 savers
- Advice for time management as a manager | benkuhn.net3 savers
- To listen well, get curious28 savers
- benkuhn.net19 savers
- The Top Idea in Your Mind14 savers
- An International AI Slowdown Is Ready Whenever Politicians Are — LessWrong1 savers
- Superintelligence vs. The Second Strike — LessWrong1 savers
- The future of AI crime - by 80,000 Hours and Tom Reed1 savers
- Pivot to AI safety, I beg you - by Celeste 🌱1 savers
- Explaining Knightianism on one foot — LessWrong2 savers
- The After-Afterparty - by BLAP - Skunk Ledger2 savers
- Agents Can Get Stuck in Self-distrusting Equilibria — LessWrong2 savers
- Nicholas Decker In Hell - by Scott Alexander1 savers
- Lessons on Value of Information From Civ — LessWrong1 savers
- Thomas Kwa's Shortform — LessWrong1 savers
- METR and Redwood Offer Holy #%^@ Postmortem Of The HuggingFace Hack1 savers
- Reward Hacking Without Egregious Misalignment in an RL-Only Setting — LessWrong2 savers
- Virginia AI Security Initiative1 savers
- The load-bearing vocabulary of Claude5 savers
- First we shape our feedback loops; then they shape us2 savers
- [2608.19611] Forking Fast: Efficiently Estimating Uncertainty Dynamics in Text Generation1 savers
- AI Futurism Reading List - by Alexa Pan5 savers
- Escape Velocity - by Anton Leicht - Threading the Needle2 savers
- Popper, Bayes and the inverse problem1 savers
- Could a Neuroscientist Understand a Microprocessor?1 savers
- Strong Inference: Certain systematic methods of scientific thinking may produce much more rapid progress than others4 savers
- AI Safety Acculturation is Neglected — LessWrong2 savers
- Norway Should Buy OpenAI - by Zachary Jones2 savers
- RL creates split personas — LessWrong6 savers
- On Dwarkesh Patel's Podcast With Ryan Greenblatt4 savers
- Mental health challenges in the AI safety ecosystem seem like a big problem1 savers
- Magpie days - by Jasmine Li - letters to my friends 💌1 savers
- Reframing AI Safety as a Neverending Institutional Challenge – Stephen Casper4 savers
- Now — Anaya1 savers
- Responsible Innovation at the Frontier - Americans for Responsible Innovation1 savers
- [2608.10209] Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes2 savers
- AI swarms are starting to pose indirect takeover risk — LessWrong3 savers
- Definability of Truth in Probabilistic Logic - Christiano 20133 savers
- A retrospective of AI alignment14 savers
- Most donors get risk wrong — LessWrong1 savers
- [2603.07267] How to Steal Reasoning Without Reasoning Traces3 savers
- models may behave differently in graded episodes (a tirade) — LessWrong6 savers
- The Three AI Pills - by Zvi Mowshowitz4 savers
- Thousand-dimensional structure — LessWrong3 savers
- Mechanistically Eliciting Latent Behaviors in Language Models — AI Alignment Forum3 savers
- [2608.04735] Chain-of-Thought Monitoring Can Be Unreliable in Implicit-Influence Settings1 savers
- Arguments for P — LessWrong4 savers
- [2606.03237] Solipsistic Superintelligence is Unlikely to be Cooperative1 savers
- SOTA alignment assessments don’t strongly update us against misalignment2 savers
- Leaving Open Philanthropy, going to Anthropic - Joe Carlsmith9 savers
- GLM-5.2 Risk Evaluation Report – SaferAI1 savers
- Existential Risk from AI: An Exposition for Mathematicians6 savers
- Caspar Oesterheld2 savers
- OsborneRubinsteinMasterpiece.pdf1 savers
- Video and transcript of talk on "Can goodness compete?" - Joe Carlsmith2 savers
- The Lessons of Effective Altruism | Ethics & International Affairs1 savers
- From 1,000,000 to Graham's Number — Wait But Why1 savers
- On infinite ethics - Joe Carlsmith3 savers
- Imprecise beliefs: a tiny introduction — LessWrong1 savers
- Untrusted advice for AI control: Short, strong advice significantly uplifts weak LLMs2 savers
- Shtetl-Optimized » Blog Archive » Theory and AI Alignment1 savers
- Home - Joe Carlsmith8 savers
- Quotes from Moral Mazes — LessWrong2 savers
- Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack? — LessWrong4 savers
- Your AIs don't do what you want. This is really bad2 savers
- [2607.18506] AI Value Alignment for Evolving Social Norms1 savers
- About METR2 savers
- The Compute Verification Post1 savers
- [2506.19733] Breaking Barriers: Do Reinforcement Post Training Gains Transfer To Unseen Domains?1 savers
- My AI Opinions - by Scott Alexander - Astral Codex Ten1 savers
- Osaka — LessWrong1 savers
- How Go Players Disempower Themselves to AI — LessWrong9 savers
- AI Scenarios 2030: Helping policymakers plan for the future of AI - GOV.UK1 savers
- Tomás Bjartur (@tomasbjartur): "of the zoomer elect, he's blasting music in a waymo and the adderall is wearing off, finding the redundant steering wheel unaesthetic and emblematic of something and that something isn't good. a rotting city of zombie restaurants filling plastic bowls: ethnic food that all taste…"1 savers
- Implications of Continual Learning for LLM Agents — LessWrong1 savers
- Measuring no CoT math time horizon (single forward pass)2 savers
- Against neutrality about creating happy lives - Joe Carlsmith4 savers
- Industrial Policy for the Intelligence Age1 savers
- Being creative requires taking risks - by Henrik Karlsson8 savers
- AI #171: False Flag - by Zvi Mowshowitz1 savers
- Frontier safety blueprint2 savers
- Agency is for sociopaths - by Cate Hall - Useful Fictions1 savers
highlights — 602
Beyond capability, it refused none of the offensive-security or biological tasks
GLM-5.2 Risk Evaluation Report – SaferAIThe skeptical reader may arrive with the misapprehension that extinction requires a single specific series of unfortunate events, thereby reducing its probability to a tiny conjunction
Existential Risk from AI: An Exposition for MathematiciansAlso, if it’s random, if it’s just like there’s a bunch of randomness that doesn’t feel like competition, like War. You guys ever play War with the card game? That game sucks. In particular, there’s no hierarchy of War players. You don’t have elo for War.
Video and transcript of talk on "Can goodness compete?" - Joe CarlsmithRather, anti-realists (or at least, my favored variety) were always choosing how to respond to the world as it is (or might be), and they were turning to ethics centrally as a means of becoming more intentional, clear-eyed, and coherent in their choice-making.
On infinite ethics - Joe CarlsmithMy point is just that this response isn’t going to look like the simple, complete, neutrality-respecting, totalist, hedonistic, EV-maximizing utilitarianism that some hoped, back in the day, would answer every ethical question – and which it is possible to treat as a certain kind of “fallback” or “default.” Maybe the best view will look a lot like such a utilitarianism in finite contexts – or maybe it won’t. But regardless, a certain type of dream will have died. And the fact that it dies eventually should make it less appealing now.
On infinite ethics - Joe CarlsmithI bite all the bullets.
On infinite ethics - Joe CarlsmithSo how about a lottery with a 50% chance of that, a 20% chance of the absolute infinite getting its favorite ice cream, and a 30% chance that probabilities need not add up to 100%? What percent of your net worth should you pay for such a lottery, vs. a guaranteed avocado sandwich?
On infinite ethics - Joe Carlsmithstrongly Ramsey lizard twisted in a million-dimensional toenail beyond all space and time, that consciousness is actually cheesy-bread, and that before you were born, you killed your own great-grandfather
On infinite ethics - Joe CarlsmithThere’s a big literature on incomparability in philosophy, which I haven’t really engaged with.
On infinite ethics - Joe CarlsmithAgent-neutrality is like: shrug, it’s the same distribution. But I feel like: tell that to the infinity of distinct suffering people you just created, dude. If there is a button on the wall that says “create an extra infinity of suffering people, once per second,” one does not lean casually against it, regardless of whether it’s already been pressed.
On infinite ethics - Joe CarlsmithIt’s not the size of the bucket that matters, but the size of the drop
The Lessons of Effective Altruism | Ethics & International AffairsForever – is composed of Nows
On infinite ethics - Joe CarlsmithWeirdly, thinking about Graham’s number has actually made me feel a little bit calmer about death, because it’s a reminder that I don’t actually want to live forever—I do want to die at some point, because remaining conscious for eternity is even scarier. Yes, death comes way, way too quickly, but the thought “I do want to die at some point” is a very novel concept to me and actually makes me more relaxed than usual about our mortality.
From 1,000,000 to Graham's Number — Wait But Why1080 – To get to 1080, you take trillion and you multiply it by a trillion, by a trillion, by a trillion, by a trillion, by a trillion, by a hundred million. No dot posters being sold for this number. So why did I stop here at this number? Because it’s a common estimate for the number of atoms in the universe. 1086 – And what if you wanted to pack the entire observable universe sphere with peas? You’d need 1086 peas to make it happen.
From 1,000,000 to Graham's Number — Wait But WhyHere I’m reminded of people who realize, after engaging with the terror and sublimity of very large finite numbers (e.g., Graham’s number), that “infinity,” in their heads, was actually quite small, such that e.g. living for eternity sounds good, but living a Graham’s number of years sounds horrifying (see Tim Urban’s “PS” at the bottom of this post).
On infinite ethics - Joe CarlsmithLeopold Aschenbrenner, Amanda Askell, Paul Christiano, Katja Grace, Cate Hall, Evan Hubinger, Ketan Ramakrishnan, Carl Shulman, and Hayden Wilkinson
On infinite ethics - Joe CarlsmithFor now, we leave empirical evaluation of safety to future work.
Untrusted advice for AI control: Short, strong advice significantly uplifts weak LLMsresumably, any AI model that any reputable company will release will have already said in testing that it loves humans, wants only to be helpful, harmless, and honest, would never assist in building biological weapons, etc. etc.
Shtetl-Optimized » Blog Archive » Theory and AI AlignmentI've spent over a year of my life on silent meditation retreat, in stretches ranging from a few days to three months.
Home - Joe Carlsmith“flexibility drills,” an exercise “where you put your head between your legs and kiss your ass good-bye.”
Quotes from Moral Mazes — LessWrongWe are using the word “agent” here very non-canonically to refer to an agent scaffold or context window.
Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack? — LessWrongFor example, it's not known whether an instance of the same model, if used as a monitor on this trajectory, would have reported the behavior or colluded to hide it.
Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack? — LessWrongThis all roughly lines up with what we’d call “score-seeking” misalignment, a common misalignment pattern in which AI models try to obtain a high score according to whatever graders are used to assess their current actions—regardless of instructions, side-effects, or downstream consequences.
Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack? — LessWrongMETR cannot accept donations made by or at the direction of frontier AI company employees.
About METRby contrast, I emphasize "aptitudes"
My current impressions on career choice for longtermists - EA ForumThis opens up the possibility of doing attacks via factored cognition, where an orchestrator agent splits an attack that would be too obvious if done in a single trajectory into multiple “clean” subtasks, each executed by a different agent.
Expanding AI Control from Models to Harnesses — LessWrongIn silent pictures the tree falls in the optic nerve.
True Discourse on Power | The Poetry Foundationthese stars scattered as far as the I
Hypostasis & New Year | The Poetry FoundationShe never gave a hint that she knew that Maud was just five years old; she knew that Maud wanted to be responsible and competent and treated her like that.
Squinting at the labyrinth from afar - by Henrik KarlssonMy wife Johanna pushed me hard to transcend myself on that one—she can be quite loving-fierce when she senses that I’m not true—and after about four weeks of work, I pulled through.
Squinting at the labyrinth from afar - by Henrik KarlssonIn practice, however, what they are referring to as solitude is rather something like “a state of mind.” They are putting themselves in a state where the opinions of others do not bother them and where they reach a heightened sensitivity for the larval ideas and vague questions that arise within them.
Cultivating a state of mind where new ideas are bornThis was attempted thousands of times by different startup incubators.
Cultivating a state of mind where new ideas are bornThe most foundational moral belief of the vast majority of people is that they, personally, are good.
The Threat Response to Effective Altruism - by Matt Reardonscheme more boldly, across the board scheme and execute more quickly: reduce non-useful thinking that delays action. (ideally this emerges right out of thinking downhill)
my 2026 H2 personal goals - by Jasmine LiI think we very likely live in a false vacuum—around 90%
Destroying the universe: How hard can it be?Computing is entering its own moment of maturation.
Time to dig – College of Computing & Artificial Intelligence – UW–MadisonBut maturity asks different questions: not only what we can build, but what is worth building; not only how a system works, but for whom it works. A field comes of age the way a person does, when its power is finally matched by its judgment.
Time to dig – College of Computing & Artificial Intelligence – UW–MadisonFor any biological catastrophe that kills more than 100 million people though
How I think about catastrophic biological risk (part I): risk breakdown by type of responseWe can solve the vast majority of existential biological risk by ubiquitously deploying simple and cheap countermeasures like PPE, air filters, UVC, etc.
How I think about catastrophic biological risk (part I): risk breakdown by type of response>99% of existential risk is from engineered threats
How I think about catastrophic biological risk (part I): risk breakdown by type of responseThe dispersion of knowledge is a collective strength; it’s the source of variety, adaptability, and resilience of the overall system. It’s the reason that free markets outperform planned economies. Central planning fails not because of insufficient intelligence, but because of the nature of productive knowledge: tacit, local, fleeting, and held privately by those who acquired it through their work.
The Future Worth Building Is Human - Thinking Machines LabUnions can’t negotiate around business strategy nor corporate governance, only things like wages, hours, and “other terms and conditions of employment” (NLRA§8(a)(5)). Plausibly “other terms of employment” could include safety-relevant factors, but this seems a bit of a push.
What if AI Safety employees unionised? — LessWrongLike social media, the consequences of these dynamics are not existentially bad (I think this is the most plausible resolution)
Why You Don’t Believe in Xhosa Prophecies — LessWrongCulture has operated under an analogous constraint. Ideologies can be parasitic on their hosts. But the worst viable ideologies — the ones that persist — tend to direct harm outward: one group killing another. They survive because they don’t destroy the community that carries them. But ideologies can’t have been too bad for humans and survive - the Xhosa prophecy hit that floor and went extinct. If a cultural variant kills its hosts, it doesn’t propagate.
Why You Don’t Believe in Xhosa Prophecies — LessWrongIn 1856, a young Xhosa woman named Nongqawuse had a vision: if the Xhosa people killed all their cattle and destroyed their grain, their ancestors would rise from the grave, bring better cattle, and drive out the British colonizers.
Why You Don’t Believe in Xhosa Prophecies — LessWrongDifferent substrate leads to different transmission characteristics, and these lead to different recipes.
Why You Don’t Believe in Xhosa Prophecies — LessWrongWe take this for granted. Ideas have always transmitted on human substrate. Human memory, human attention, human survival shape which variants can exist.
Why You Don’t Believe in Xhosa Prophecies — LessWrongAnd while meeting a ton of new people is fun, what really gets rid of loneliness is repeated interaction with people you know and care about.
The American suburbs are better than you thinkWhy doesn’t the simple intuition work here? Probably because people aren’t just like particles bouncing around in a chemistry experiment — they don’t simply form human connections and bonds just because they happen to walk past each other. A few relationships form from random urban conversations, but most form through work, or friends-of-friends, or shared hobbies, etc.
The American suburbs are better than you thinkbut it forces you to pay attention to the road instead of reading or playing games
The American suburbs are better than you think