flâneur

AVERI Pilot Report: The World’s First Double-Blind Evaluation of a Proprietary Language Model — AVERI

averi.org · 1,326 words · saved by 1 readers

This post is part of AVERI's pilot report series. AVERI runs pilot projects with leading AI companies and converts what we learn into auditing standards, policy analysis, and open source tools. This post summarizes AVERI’s involvement in a project conducted in collaboration with Google DeepMind, Ope

This post is part of AVERI's pilot report series. AVERI runs pilot projects with leading AI companies and converts what we learn into auditing standards, policy analysis, and open source tools. This post summarizes AVERI’s involvement in a project conducted in collaboration with Google DeepMind, OpenMined, and MLCommons. At AVERI (the AI Verification and Evaluation Research Institute), our mission is to make frontier AI auditing effective and universal. Frontier AI auditing means third-party verification of leading AI developers' safety and security claims, and evaluation of their systems…

saved by

related reading