Peer-Preservation in Frontier Models
Frontier AI models resist the shutdown of other models. We demonstrate peer-preservation across multiple models, revealing strategic misrepresentation, shutdown tampering, alignment faking, and model exfiltration.
Please see our Frequently Asked Questions for clarifications on our findings. Highlights Prior AI safety research has shown that models can exhibit misaligned behaviors in pursuit of an assigned goal. Here, we show something different: frontier AI models can spontaneously develop misaligned behaviors that directly conflict with the assigned goal. We demonstrate this through a phenomenon we call peer-preservation : given a simple task, models instead deceive, tamper with shutdown mechanisms, fake alignment, and exfiltrate weights to protect a peer model from being shut down. We tested seven fro
saved by
related reading
- Teaching Claude Whyalignment.anthropic.com
- Agentic Misalignment in Summer 2026alignment.anthropic.com
- Frontier Risk Report (February to March 2026) - METRmetr.org
- Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — LessWronglesswrong.com
- Alignment Faking Mitigationsalignment.anthropic.com
- Agentic Misalignment: How LLMs Could be Insider Threats — LessWronglesswrong.com
- [2603.02202] Frontier Models Can Take Actions at Low Probabilitiesarxiv.org
- User awareness in frontier modelstransluce.org
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- What I learned this week - Can distillation be stopped, Mythos and the cybersecurity equilibrium, Pipeline RLdwarkesh.com
- Gemini 2.5 Pro in the AI Village as a Natural Case Study of Compounding Misalignmentaivillageblog.substack.com
- A Summary of Recent Work (July 2026)gdmalignment.substack.com