✳flâneur — a map of the web's best reading
The Most Forbidden Technique — LessWrong
lesswrong.com · 5,935 words · saved by 1 readers
The Most Forbidden Technique is training an AI using interpretability techniques. …
x The Most Forbidden Technique — LessWrong Orthogonality Thesis AI Frontpage 2025 Top Fifty: 27 % 167 The Most Forbidden Technique by Zvi 12th Mar 2025 Don't Worry About the Vase 20 min read 9 167 The Most Forbidden Technique is training an AI using interpretability techniques. An AI produces a final output [X] via some method [M]. You can analyze [M] using technique [T], to learn what the AI is up to. You could train on that. Never do that. You train on [X]. Only [X]. Never [M], never [T]. Why? Because [T] is how you figure out when the model is misbehaving. If you train on [T], you are train
Explore this link on the map →saved by
related reading
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Neel Nanda on Mechanistic Interpretability: Progress, Limits, and Paths to Safer AI — EA Forumforum.effectivealtruism.org
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscationarxiv.org
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- How confessions can keep language models honest | OpenAIopenai.com
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com
- the case for CoT unfaithfulness is overstated — LessWronglesswrong.com
- When Chain of Thought is Necessary, Language Models Struggle to Evade Monitorsarxiv.org
- Research Areas in Methods for Post-training and Elicitation (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- The fragile foundations of CoT monitoring | Christopher Pottsweb.stanford.edu
- Paper: Prompt Optimization Makes Misalignment Legible — LessWronglesswrong.com