ML Systems Will Have Weird Failure Modes
Previously, I've argued that future ML systems might exhibit unfamiliar, emergent capabilities [https://bounded-regret.ghost.io/p/1527e9dd-c48d-4941-9b14-4f7293318d5c/], and that thought experiments provide one approach [https://bounded-regret.ghost.io/p/a2d733a7-108a-4587-97fb-db90f66ce030/] towards predicting these capabilities and their consequences. In this post I’ll describe a particular thought experiment in detail.
Previously, I've argued that future ML systems might exhibit unfamiliar, emergent capabilities , and that thought experiments provide one approach towards predicting these capabilities and their consequences. In this post I’ll describe a particular thought experiment in detail. We’ll see that taking thought experiments seriously often surfaces future risks that seem "weird" and alien from the point of view of current systems. I’ll also describe how I tend to engage with these thought experiments: I usually start out intuitively skeptical, but when I reflect on emergent behavior I find that som
Explore this link on the map →related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- 39 - Evan Hubinger on Model Organisms of Misalignment | AXRP - the AI X-risk Research Podcastaxrp.net
- Emergent Deception and Emergent Optimizationbounded-regret.ghost.io
- Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forumalignmentforum.org
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Thought Experiments Provide a Third Anchorbounded-regret.ghost.io
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- How likely is deceptive alignment? — AI Alignment Forumalignmentforum.org
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- Research Areas in Evaluation and Guarantees in Reinforcement Learning (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- Mediumai-alignment.com