Peer-Preservation in Frontier Models
Frontier AI models resist the shutdown of other models. We demonstrate peer-preservation across multiple models, revealing strategic misrepresentation, shutdown tampering, alignment faking, and model exfiltration.
Please see our Frequently Asked Questions for clarifications on our findings. Highlights Prior AI safety research has shown that models can exhibit misaligned behaviors in pursuit of an assigned goal. Here, we show something different: frontier AI models can spontaneously develop misaligned behaviors that directly conflict with the assigned goal. We demonstrate this through a phenomenon we call peer-preservation : given a simple task, models instead deceive, tamper with shutdown mechanisms, fake alignment, and exfiltrate weights to protect a peer model from being shut down. We tested seven fro
Explore this link on the map →saved by
related reading
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Agentic Misalignment: How LLMs Could be Insider Threats — LessWronglesswrong.com
- Frontier Risk Report (February to March 2026) - METRmetr.org
- [2603.02202] Frontier Models Can Take Actions at Low Probabilitiesarxiv.org
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- What I learned this week - Can distillation be stopped, Mythos and the cybersecurity equilibrium, Pipeline RLdwarkesh.com
- Agentic Misalignment in Summer 2026alignment.anthropic.com
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forumalignmentforum.org
- 39 - Evan Hubinger on Model Organisms of Misalignment | AXRP - the AI X-risk Research Podcastaxrp.net