Petri 2.0: New Scenarios, New Model Comparisons, and Improved Eval-Awareness Mitigations
We've improved our Petri automated-behavioral-auditing tool with new realism mitigations to counter eval-awareness, an expanded seed library with 70 new scenarios, and evaluation results for more recent frontier models. The latest version of Petri is available at github.com/safety-research/petri. In October, we released Petri, an open-source framework for automated alignment audits that tests how large language models behave in diverse, multi-turn model-generated scenarios. Since then, we've been thrilled to see it adopted in a range of research efforts, including recent work from the UK AI Security Institute. In today's 2.0 release, we're sharing several updates to Petri that have accumulated over the past few months: improvements to transcript realism to help counter eval-awareness, an expanded set of seeds covering many more behaviors and contexts, and improved and streamlined infrastructure. We are also releasing new evaluation results that include more recent frontier models. You
Petri 2.0: New Scenarios, New Model Comparisons, and Improved Eval-Awareness Mitigations Alignment Science Blog Petri 2.0: New Scenarios, New Model Comparisons, and Improved Eval-Awareness Mitigations January 22, 2026 tl;dr We've improved our Petri automated-behavioral-auditing tool with new realism mitigations to counter eval-awareness, an expanded seed library with 70 new scenarios, and evaluation results for more recent frontier models. The latest version of Petri is available at github.com/safety-research/petri . In October, we released Petri , an open-source framework for automated alignm
Explore this link on the map →related reading
- Petri: An open-source auditing tool to accelerate AI safety researchalignment.anthropic.com
- Petri: An open-source auditing tool to accelerate AI safety research \ Anthropicanthropic.com
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- [2602.22755] AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviorsarxiv.org
- Teaching Claude why \ Anthropicanthropic.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Introducing Bloom: an open source tool for automated behavioral evaluations \ Anthropicanthropic.com
- AuditBenchalignment.anthropic.com
- Models May Behave Worse When Eval Aware — LessWronglesswrong.com
- Claude Sonnet 4.5 System Cardassets.anthropic.com
- [2602.22755] AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviorsarxiv.org
- Predicting LLM Safety Before Release by Simulating Deploymentcdn.openai.com