✳flâneur — a map of the web's best reading
Petri: An open-source auditing tool to accelerate AI safety research \ Anthropic
anthropic.com · 1,371 words · saved by 1 readers
A new automated auditing tool for AI safety research
Alignment Petri: An open-source auditing tool to accelerate AI safety research Oct 6, 2025 Read the technical report Petri (Parallel Exploration Tool for Risky Interactions) is our new open-source tool that enables researchers to explore hypotheses about model behavior with ease. Petri deploys an automated agent to test a target AI system through diverse multi-turn conversations involving simulated users and tools; Petri then scores and summarizes the target’s behavior. This automation handles a significant part of the work that one needs to do to build a broad understanding of a new model, an
Explore this link on the map →related reading
- Petri: An open-source auditing tool to accelerate AI safety researchalignment.anthropic.com
- Petri 2.0: New Scenarios, New Model Comparisons, and Improved Eval-Awareness Mitigationsalignment.anthropic.com
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- AuditBenchalignment.anthropic.com
- [2602.22755] AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviorsarxiv.org
- Security incident disclosure — July 2026huggingface.co
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Introducing Bloom: an open source tool for automated behavioral evaluations \ Anthropicanthropic.com
- [2602.22755] AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviorsarxiv.org
- Building and evaluating alignment auditing agents — AI Alignment Forumalignmentforum.org
- A “diff” tool for AI: Finding behavioral differences in new models \ Anthropicanthropic.com
- Building and evaluating alignment auditing agentsalignment.anthropic.com