Your Model Organisms Might Be Fried — LessWrong
Context: We are the ‘model motivations’ team at Arcadia Alignment. We aim to build a science of ‘model intentions’, unifying insights from personas a…
x Your Model Organisms Might Be Fried — LessWrong AI Frontpage 81 Your Model Organisms Might Be Fried by Daniel Tan , J Bostock , draganover , ma-rmartinez , sidbaines , David Africa 18th Jun 2026 8 min read 5 81 Context: We are the ‘model motivations’ team at Arcadia Alignment. We aim to build a science of ‘ model intentions ’, unifying insights from personas and other empirical evidence. This is an informal research note that has come out of the first 2-3 weeks of exploratory work. In this post, we’ll outline the need for much better model organisms and how we might get there. The case for b
Explore this link on the map →saved by
related reading
- Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forumalignmentforum.org
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forumalignmentforum.org
- 39 - Evan Hubinger on Model Organisms of Misalignment | AXRP - the AI X-risk Research Podcastaxrp.net
- Model Organisms for Emergent Misalignment — LessWronglesswrong.com
- Advice for making robust-to-training model organismsblog.redwoodresearch.org
- From personas to intentions: towards a science of motivations for AI models — LessWronglesswrong.com
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Model organisms researchers should check whether high LRs defeat their model organisms — LessWronglesswrong.com
- Auditing language models for hidden objectives — LessWronglesswrong.com
- Alignment faking in large language modelsarxiv.org