TASTE: Can AI Models Judge AI Safety Research Proposals?
We built TASTE (The AI Safety Taste Evaluation) — a benchmark measuring how well models can judge pairs of AI safety research proposals, scored by agreement with the preferences of experienced human researchers. Two design choices were important for building a high-agreement benchmark (92 pairs, 77% estimated human agreement): a discussion stage in which researchers talk through disagreements before revising their scores, and filtering researchers’ labels for self-reported "strong" confidence. We find models perform worse than human researchers on TASTE (Fable 5, 60%).
Hasan Baig1, Hailey Joren2, Joe Benton2 August 28, 2026 1Anthropic Fellows Program; 2Anthropic tl;dr We built TASTE (The AI Safety Taste Evaluation) — a benchmark measuring how well models can judge pairs of AI safety research proposals, scored by agreement with the preferences of experienced human researchers. Two design choices were important for building a high-agreement benchmark (92 pairs, 77% estimated human agreement): a discussion stage in which researchers talk through disagreements before revising their scores, and filtering researchers’ labels for self-reported "strong"…
saved by
related reading
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- CAIS AI Dashboarddashboard.safe.ai
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- The Universe from an Intentional Stancecasparoesterheld.com
- Toward A Public Science of Model Behavior | Transluce AItransluce.org
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Predicting Empirical AI Research Outcomes with Language Modelsarxiv.org
- Research Areas in Benchmark Design and Evaluation (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- AGI safety career advice — EA Forumforum.effectivealtruism.org
- In pursuit of a benchmark for human tastenotes.designarena.ai
- Challenges in evaluating AI systems \ Anthropicanthropic.com