Advice for making robust-to-training model organisms
We’d like to develop training techniques that work when applied to future misaligned AI systems. One strategy for studying proposed techniques is to test them on model organisms. However, model organisms built with common techniques are often fragile: we (and other researchers like
Advice for making robust-to-training model organisms +1 Alek Westover , Sebastian Prasanna , Vivek Hebbar , and 2 others May 28, 2026 12 Share We’d like to develop training techniques that work when applied to future misaligned AI systems . One strategy for studying proposed techniques is to test them on model organisms . However, model organisms built with common techniques are often fragile: we (and other researchers like Roger et al. and Ryd et al. ) have observed them to stop misbehaving after untargeted training—training that doesn’t directly target the misbehavior. For example, we have o
Explore this link on the map →saved by
related reading
- Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forumalignmentforum.org
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Your Model Organisms Might Be Fried — LessWronglesswrong.com
- [2401.05566] Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Trainingarxiv.org
- 39 - Evan Hubinger on Model Organisms of Misalignment | AXRP - the AI X-risk Research Podcastaxrp.net
- [2604.22082] Removing Sandbagging in LLMs by Training with Weak Supervisionarxiv.org
- A small number of samples can poison LLMs of any size \ Anthropicanthropic.com
- Teaching Claude why \ Anthropicanthropic.com
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- Model organisms researchers should check whether high LRs defeat their model organisms — LessWronglesswrong.com
- Model Organisms for Emergent Misalignment — LessWronglesswrong.com
- You Can Catch Sleeper Agents by Teaching Another Model to Imitate Them — LessWronglesswrong.com