Joshua Achiam on X: "A very hot take: chain of thought interpretability was always going to be so fragile as to be an unacceptable backstop for long-term AI safety, and while I admire the optimism and effort involved in protecting its fidelity (and consider such effort to have been worthwhile), I do" / X
A very hot take: chain of thought interpretability was always going to be so fragile as to be an unacceptable backstop for long-term AI safety, and while I admire the optimism and effort involved in protecting its fidelity (and consider such effort to have been worthwhile), I do not think it makes …
A very hot take: chain of thought interpretability was always going to be so fragile as to be an unacceptable backstop for long-term AI safety, and while I admire the optimism and effort involved in protecting its fidelity (and consider such effort to have been worthwhile), I do not think it makes sense to elevate as a principle the idea that the chain of thought must remain legible to humans. I would go so far as to say that strategies predicated on that principle are definitely doomed, in that they will not work eventually, and we should not depend on them or take enduring reassurance from…
saved by
related reading
- Dario Amodei — The Urgency of Interpretabilitydarioamodei.com
- A Summary of Recent Work (July 2026)gdmalignment.substack.com
- Neel Nanda on Mechanistic Interpretability: Progress, Limits, and Paths to Safer AI — EA Forumforum.effectivealtruism.org
- [2507.11473] Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safetyarxiv.org
- AI in 2025: gestalt — LessWronglesswrong.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- The fragile foundations of CoT monitoring | Christopher Pottsweb.stanford.edu
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- How AI Is Learning to Think in Secretnickandresen.substack.com
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- How Can Interpretability Researchers Help AGI Go Well? — AI Alignment Forumalignmentforum.org
- AI #24: Week of the Podcast — LessWronglesswrong.com