Petri 2.0: New Scenarios, New Model Comparisons, and Improved Eval-Awareness Mitigations
We've improved our Petri automated-behavioral-auditing tool with new realism mitigations to counter eval-awareness, an expanded seed library with 70 new scenarios, and evaluation results for more recent frontier models. The latest version of Petri is available at github.com/safety-research/petri. In October, we released Petri, an open-source framework for automated alignment audits that tests how large language models behave in diverse, multi-turn model-generated scenarios. Since then, we've been thrilled to see it adopted in a range of research efforts, including recent work from the UK AI Security Institute. In today's 2.0 release, we're sharing several updates to Petri that have accumulated over the past few months: improvements to transcript realism to help counter eval-awareness, an expanded set of seeds covering many more behaviors and contexts, and improved and streamlined infrastructure. We are also releasing new evaluation results that include more recent frontier models. You
Petri 2.0: New Scenarios, New Model Comparisons, and Improved Eval-Awareness Mitigations Alignment Science Blog Petri 2.0: New Scenarios, New Model Comparisons, and Improved Eval-Awareness Mitigations January 22, 2026 tl;dr We've improved our Petri automated-behavioral-auditing tool with new realism mitigations to counter eval-awareness, an expanded seed library with 70 new scenarios, and evaluation results for more recent frontier models. The latest version of Petri is available at github.com/safety-research/petri . In October, we released Petri , an open-source framework for automated alignm
Explore this link on the map →related reading
- Petri: An open-source auditing tool to accelerate AI safety researchalignment.anthropic.com
- Petri: An open-source auditing tool to accelerate AI safety research \ Anthropicanthropic.com
- Teaching Claude Whyalignment.anthropic.com
- An alignment assessment of recent cybersecurity incidentsanthropic.com
- Sidestepping Evaluation Awareness and Anticipating Misalignment with Production Evaluationsalignment.openai.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Agentic Misalignment in Summer 2026alignment.anthropic.com
- Alignment Faking Mitigationsalignment.anthropic.com
- Claude Sonnet 4.5 System Cardassets.anthropic.com
- [2602.22755] AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviorsarxiv.org
- Toward A Public Science of Model Behavior | Transluce AItransluce.org
- [2605.24229] How Well Do Models Follow Their Constitutions?arxiv.org