Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forum
TL;DR: This document lays out the case for research on “model organisms of misalignment” – in vitro demonstrations of the kinds of failures that might pose existential threats – as a new and important pillar of alignment research. If you’re interested in working on this agenda with us at Anthropic, we’re hiring! Please apply to the research scientist or research engineer position on the Anthropic website and mention that you’re interested in working on model organisms of misalignment. We don’t currently have ~any strong empirical evidence for the most concerning sources of existential risk, most notably stories around dishonest AI systems that actively trick or fool their training processes or human operators: A significant part of why we think we don't see empirical examples of these failure modes is that they require that the AI develop several scary tendencies and capabilities such as situational awareness and deceptive reasoning and deploy them together. As a result, it seems usefu
x Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forum Best of LessWrong 2023 AI Evaluations Deceptive Alignment Deception Language Models (LLMs) Research Agendas AI Curated 124 Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research by evhub , Nicholas Schiefer , Carson Denison , Ethan Perez 8th Aug 2023 22 min read 30 124 TL;DR : This document lays out the case for research on “model organisms of misalignment” – in vitro demonstrations of the kinds of failures that might pose existential threats – as a new and importan
Explore this link on the map →related reading
- Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forumalignmentforum.org
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- 39 - Evan Hubinger on Model Organisms of Misalignment | AXRP - the AI X-risk Research Podcastaxrp.net
- Your Model Organisms Might Be Fried — LessWronglesswrong.com
- Model Organisms for Emergent Misalignment — LessWronglesswrong.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Research Areas in Evaluation and Guarantees in Reinforcement Learning (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Teaching Claude why \ Anthropicanthropic.com
- Alignment remains a hard, unsolved problem — AI Alignment Forumalignmentforum.org
- Chapter 4: Alignment Science - ARENAlearn.arena.education
- Externalized reasoning oversight: a research direction for language model alignment — AI Alignment Forumalignmentforum.org