Petri: An open-source auditing tool to accelerate AI safety research \ Anthropic
anthropic.com · 1,371 words · saved by 1 readers
A new automated auditing tool for AI safety research
Alignment Petri: An open-source auditing tool to accelerate AI safety research Oct 6, 2025 Read the technical report Petri (Parallel Exploration Tool for Risky Interactions) is our new open-source tool that enables researchers to explore hypotheses about model behavior with ease. Petri deploys an automated agent to test a target AI system through diverse multi-turn conversations involving simulated users and tools; Petri then scores and summarizes the target’s behavior. This automation handles a significant part of the work that one needs to do to build a broad understanding of a new model, an
related reading
- Petri: An open-source auditing tool to accelerate AI safety researchalignment.anthropic.com
- Petri 2.0: New Scenarios, New Model Comparisons, and Improved Eval-Awareness Mitigationsalignment.anthropic.com
- Pre-deployment auditing can catch an overt saboteuralignment.anthropic.com
- Security incident disclosure — July 2026huggingface.co
- Toward A Public Science of Model Behavior | Transluce AItransluce.org
- An alignment assessment of recent cybersecurity incidentsanthropic.com
- [2602.22755] AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviorsarxiv.org
- A “diff” tool for AI: Finding behavioral differences in new models \ Anthropicanthropic.com
- Frontier Risk Report (February to March 2026) - METRmetr.org
- Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — LessWronglesswrong.com
- Privacy-Preserving AI Audit Tools — OpenMinedopenmined.org
- Auditing language models for hidden objectives — LessWronglesswrong.com