Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forum
TL;DR: This document lays out the case for research on “model organisms of misalignment” – in vitro demonstrations of the kinds of failures that migh…
x Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forum Best of LessWrong 2023 AI Evaluations Deceptive Alignment Deception Language Models (LLMs) Research Agendas AI Curated 123 Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research by evhub , Nicholas Schiefer , Carson Denison , Ethan Perez 8th Aug 2023 22 min read 30 123 TL;DR : This document lays out the case for research on “model organisms of misalignment” – in vitro demonstrations of the kinds of failures that might pose existential threats – as a new and importan
saved by
related reading
- Model Organisms for Emergent Misalignment — LessWronglesswrong.com
- Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forumalignmentforum.org
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- 39 - Evan Hubinger on Model Organisms of Misalignment | AXRP - the AI X-risk Research Podcastaxrp.net
- Alignment is not solved but it increasingly looks solvablealigned.substack.com
- Auditing language models for hidden objectives — LessWronglesswrong.com
- New report: “Scheming AIs: Will AIs fake alignment during training in order to get power?”joecarlsmith.com
- Many arguments for AI x-risk are wrong — AI Alignment Forumalignmentforum.org
- Your Model Organisms Might Be Fried — LessWronglesswrong.com
- Alignment Faking Mitigationsalignment.anthropic.com
- Teaching Claude Whyalignment.anthropic.com
- Research Areas in Evaluation and Guarantees in Reinforcement Learning (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org