Petri: An open-source auditing tool to accelerate AI safety research
We're releasing Petri (Parallel Exploration Tool for Risky Interactions), an open-source framework for automated auditing that uses AI agents to test the behaviors of target models across diverse scenarios. When applied to 14 frontier models with 111 seed instructions, Petri successfully elicited a broad set of misaligned behaviors including autonomous deception, oversight subversion, whistleblowing, and cooperation with human misuse. The tool is available now at github.com/safety-research/petri. AI models are becoming more capable and are being deployed with wide-ranging affordances across more domains, increasing the surface area where misaligned behaviors might emerge. The sheer volume and complexity of potential behaviors far exceeds what researchers can manually test, making it increasingly difficult to properly audit each model. Over the past year, we've been building automated auditing agents to help address this challenge. We used them in the Claude 4 and Claude Sonnet 4.5 Syst
Petri: An open-source auditing tool to accelerate AI safety research Alignment Science Blog Petri: An open-source auditing tool to accelerate AI safety research October 6, 2025 tl;dr We're releasing Petri (Parallel Exploration Tool for Risky Interactions), an open-source framework for automated auditing that uses AI agents to test the behaviors of target models across diverse scenarios. When applied to 14 frontier models with 111 seed instructions, Petri successfully elicited a broad set of misaligned behaviors including autonomous deception, oversight subversion, whistleblowing, and cooperati
Explore this link on the map →related reading
- Petri: An open-source auditing tool to accelerate AI safety research \ Anthropicanthropic.com
- Petri 2.0: New Scenarios, New Model Comparisons, and Improved Eval-Awareness Mitigationsalignment.anthropic.com
- [2602.22755] AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviorsarxiv.org
- Building and evaluating alignment auditing agents — AI Alignment Forumalignmentforum.org
- AuditBenchalignment.anthropic.com
- Security incident disclosure — July 2026huggingface.co
- Auditing language models for hidden objectives — LessWronglesswrong.com
- [2602.22755] AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviorsarxiv.org
- [2604.16812] Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- Building and evaluating alignment auditing agentsalignment.anthropic.com
- 2312.06942arxiv.org