flâneur — a map of the web's best reading

Petri 2.0: New Scenarios, New Model Comparisons, and Improved Eval-Awareness Mitigations

alignment.anthropic.com · 2,380 words · saved by 1 readers

We've improved our Petri automated-behavioral-auditing tool with new realism mitigations to counter eval-awareness, an expanded seed library with 70 new scenarios, and evaluation results for more recent frontier models. The latest version of Petri is available at github.com/safety-research/petri. In October, we released Petri, an open-source framework for automated alignment audits that tests how large language models behave in diverse, multi-turn model-generated scenarios. Since then, we've been thrilled to see it adopted in a range of research efforts, including recent work from the UK AI Security Institute. In today's 2.0 release, we're sharing several updates to Petri that have accumulated over the past few months: improvements to transcript realism to help counter eval-awareness, an expanded set of seeds covering many more behaviors and contexts, and improved and streamlined infrastructure. We are also releasing new evaluation results that include more recent frontier models. You

Petri 2.0: New Scenarios, New Model Comparisons, and Improved Eval-Awareness Mitigations Alignment Science Blog Petri 2.0: New Scenarios, New Model Comparisons, and Improved Eval-Awareness Mitigations January 22, 2026 tl;dr We've improved our Petri automated-behavioral-auditing tool with new realism mitigations to counter eval-awareness, an expanded seed library with 70 new scenarios, and evaluation results for more recent frontier models. The latest version of Petri is available at github.com/safety-research/petri . In October, we released Petri , an open-source framework for automated alignm

Explore this link on the map →

related reading