✳flâneur — a map of the web's best reading
Introducing Bloom: an open source tool for automated behavioral evaluations \ Anthropic
anthropic.com · 1,401 words · saved by 1 readers
Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.
Alignment Introducing Bloom: an open source tool for automated behavioral evaluations Dec 19, 2025 Read the technical report We're releasing Bloom, an open source agentic framework for generating behavioral evaluations of frontier AI models. Bloom takes a researcher-specified behavior and quantifies its frequency and severity across automatically generated scenarios. Bloom's evaluations correlate strongly with our hand-labeled judgments and we find they reliably separate baseline models from intentionally misaligned ones. As examples of this, we release benchmark results for four alignment rel
Explore this link on the map →related reading
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- Demystifying evals for AI agents \ Anthropicanthropic.com
- Claude 4 System Cardwww-cdn.anthropic.com
- Teaching Claude why \ Anthropicanthropic.com
- Claude Sonnet 4.5 System Cardassets.anthropic.com
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- Claude Sonnet 3.7 (often) knows when it’s in alignment evaluations — LessWronglesswrong.com
- Challenges in evaluating AI systems \ Anthropicanthropic.com
- Claude Sonnet 3.7 (often) knows when it’s in alignment evaluations – Apollo Researchapolloresearch.ai
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Petri 2.0: New Scenarios, New Model Comparisons, and Improved Eval-Awareness Mitigationsalignment.anthropic.com
- Anthropic’s Transparency Hub \ Anthropicanthropic.com