double-blind-evaluations-technical-report.pdf
storage.googleapis.com · 5,088 words · saved by 1 readers
N/A
Double Blind Evals: Resolving the Dual Confidentiality Dilemma in AI Safety Auditing Andrew Trask2,4 Sol Messing 2 Vai Pahwa 2 Pegah Maham 2 Rene Kolga 2 Austin Frantz 2 Andrew Tash 2 Kate Thomas 2 1 1 1 Sean McGregor George Balston Patricia Paskov Miles Brundage 1 Akriti Vij 3 Bennett Hillenbrand 5 Alexandros Karargyris 5 2 2 Tin Acosta Joel Fenster Madeleine Eilish 2…
saved by
related reading
- AVERI Pilot Report: The World’s First Double-Blind Evaluation of a Proprietary Language Model — AVERIaveri.org
- Privacy-Preserving AI Audit Tools — OpenMinedopenmined.org
- A Safe Path to Open Weights - Thinking Machines Labthinkingmachines.ai
- Private Cloud Compute: A new frontier for AI privacy in the cloud - Apple Security Researchsecurity.apple.com
- Challenges in evaluating AI systems \ Anthropicanthropic.com
- 2312.06942arxiv.org
- Toward A Public Science of Model Behavior | Transluce AItransluce.org
- Confidential AI From GPU Enclavesblog.blyss.dev
- We need 3rd party Training-Run Assessments — LessWronglesswrong.com
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — LessWronglesswrong.com
- Claude Sonnet 4.5 System Cardassets.anthropic.com