A safe harbor for AI evaluation and red teaming
normaltech.ai · 1,795 words · saved by 1 readers
An argument for legal and technical safe harbors for AI safety and trustworthiness research
This blog post is authored by Shayne Longpre, Sayash Kapoor, Kevin Klyman, Ashwin Ramaswami, Rishi Bommasani, Arvind Narayanan, Percy Liang, and Peter Henderson. The paper has 23 authors and is available here. Today, we are releasing an open letter encouraging AI companies to provide legal and technical protections for good-faith research on their AI models. The letter focuses on the importance of independent evaluations of proprietary generative AI models, particularly those with millions of users. In an accompanying paper, we discuss existing challenges to independent research and how a…
related reading
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- A Safe Path to Open Weights - Thinking Machines Labthinkingmachines.ai
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- Challenges in evaluating AI systems \ Anthropicanthropic.com
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com
- AI safety is not a model propertysubstack.com
- AVERI Pilot Report: The World’s First Double-Blind Evaluation of a Proprietary Language Model — AVERIaveri.org
- Toward A Public Science of Model Behavior | Transluce AItransluce.org
- [2408.02565] Reasons to Doubt the Impact of AI Risk Evaluationsarxiv.org
- We need 3rd party Training-Run Assessments — LessWronglesswrong.com
- Making sense of the UK’s AI Security Instituteadalovelaceinstitute.org
- 2312.06942arxiv.org