The Most Forbidden Technique — LessWrong
lesswrong.com · 5,935 words · saved by 3 readers
The Most Forbidden Technique is training an AI using interpretability techniques. …
x The Most Forbidden Technique — LessWrong Orthogonality Thesis AI Frontpage 2025 Top Fifty: 27 % 167 The Most Forbidden Technique by Zvi 12th Mar 2025 Don't Worry About the Vase 20 min read 9 167 The Most Forbidden Technique is training an AI using interpretability techniques. An AI produces a final output [X] via some method [M]. You can analyze [M] using technique [T], to learn what the AI is up to. You could train on that. Never do that. You train on [X]. Only [X]. Never [M], never [T]. Why? Because [T] is how you figure out when the model is misbehaving. If you train on [T], you are train
saved by
related reading
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- The fragile foundations of CoT monitoring | Christopher Pottsweb.stanford.edu
- How AI Is Learning to Think in Secretnickandresen.substack.com
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscationarxiv.org
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscationarxiv.org
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- CoT_Monitoring.pdfcdn.openai.com
- Why We Are Excited About Confessionsalignment.openai.com
- [2507.11473] Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safetyarxiv.org
- Training a Misaligned Reward Seekeralignment.anthropic.com
- How confessions can keep language models honest | OpenAIopenai.com