AVERI Pilot Report: The World’s First Double-Blind Evaluation of a Proprietary Language Model — AVERI
This post is part of AVERI's pilot report series. AVERI runs pilot projects with leading AI companies and converts what we learn into auditing standards, policy analysis, and open source tools. This post summarizes AVERI’s involvement in a project conducted in collaboration with Google DeepMind, Ope
This post is part of AVERI's pilot report series. AVERI runs pilot projects with leading AI companies and converts what we learn into auditing standards, policy analysis, and open source tools. This post summarizes AVERI’s involvement in a project conducted in collaboration with Google DeepMind, OpenMined, and MLCommons. At AVERI (the AI Verification and Evaluation Research Institute), our mission is to make frontier AI auditing effective and universal. Frontier AI auditing means third-party verification of leading AI developers' safety and security claims, and evaluation of their systems…
saved by
related reading
- double-blind-evaluations-technical-report.pdfstorage.googleapis.com
- Privacy-Preserving AI Audit Tools — OpenMinedopenmined.org
- Summary of METR's predeployment evaluation of GPT-5.6 Solmetr.org
- gpt-4.pdfcdn.openai.com
- Challenges in evaluating AI systems \ Anthropicanthropic.com
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- Gemini 3 is Evaluation-Paranoid and Contaminated — LessWronglesswrong.com
- Auditing language models for hidden objectives — LessWronglesswrong.com
- GLM-5.2 Risk Evaluation Report – SaferAIsafer-ai.org
- We need 3rd party Training-Run Assessments — LessWronglesswrong.com
- Claude Sonnet 4.5 System Cardassets.anthropic.com
- metr.org/risk-report-feb-mar-2026.pdf#page=33.47metr.org