2410.21514
arxiv.org · 8,141 words · saved by 1 readers
N/A
Sabotage Evaluations for Frontier Models Joe Benton, Misha Wagner, Eric Christiansen, Cem Anil, Ethan Perez, Jai Srivastav, Esin Durmus, Deep Ganguli, Shauna Kravec, Buck Shlegeris*, Jared Kaplan, Holden Karnofsky† , Evan Hubinger, Roger Grosse, arXiv:2410.21514v1 [cs.LG] 28 Oct 2024 Samuel R. Bowman, David Duvenaud…
related reading
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- Sabotage evaluations for frontier models \ Anthropicanthropic.com
- Pre-deployment auditing can catch an overt saboteuralignment.anthropic.com
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Claude Sonnet 4.5 System Cardassets.anthropic.com
- Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — LessWronglesswrong.com
- Redacted Risk Report Feb 2026www-cdn.anthropic.com
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- [2403.13793] Evaluating Frontier Models for Dangerous Capabilitiesarxiv.org
- [2603.02202] Frontier Models Can Take Actions at Low Probabilitiesarxiv.org
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org