✳flâneur — a map of the web's best reading
How to Design Environments for Understanding Model Motives — LessWrong
lesswrong.com · 6,238 words · saved by 1 readers
Authors: Gerson Kroiz*, Aditya Singh*, Senthooran Rajamanoharan, Neel Nanda …
x How to Design Environments for Understanding Model Motives — LessWrong Interpretability (ML & AI) MATS Program AI Frontpage 51 How to Design Environments for Understanding Model Motives by gersonkroiz , aditya singh , Senthooran Rajamanoharan , Neel Nanda 2nd Mar 2026 AI Alignment Forum 12 min read 0 51 Ω 21 Authors: Gerson Kroiz*, Aditya Singh*, Senthooran Rajamanoharan, Neel Nanda Gerson and Aditya are co-first authors. This work was conducted during MATS 9.0 and was advised by Senthooran Rajamanoharan and Neel Nanda. TL;DR Understanding why a model took an action is a key question in AI S
Explore this link on the map →related reading
- How confessions can keep language models honest | OpenAIopenai.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Models May Behave Worse When Eval Aware — LessWronglesswrong.com
- confessions_paper.pdfcdn.openai.com
- Sandbagging with misaligned action - Chain-of-Thought Transcript - Anti-Schemingantischeming.ai
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- From personas to intentions: towards a science of motivations for AI models — LessWronglesswrong.com
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscationarxiv.org
- When Chain of Thought is Necessary, Language Models Struggle to Evade Monitorsarxiv.org
- A Toy Environment For Exploring Reasoning About Reward — LessWronglesswrong.com
- How well do models follow their constitutions? — LessWronglesswrong.com
- 39 - Evan Hubinger on Model Organisms of Misalignment | AXRP - the AI X-risk Research Podcastaxrp.net