Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forum
TL;DR: This document lays out the case for research on “model organisms of misalignment” – in vitro demonstrations of the kinds of failures that migh…
x Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forum Best of LessWrong 2023 AI Evaluations Deceptive Alignment Deception Language Models (LLMs) Research Agendas AI Curated 123 Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research by evhub , Nicholas Schiefer , Carson Denison , Ethan Perez 8th Aug 2023 22 min read 30 123 TL;DR : This document lays out the case for research on “model organisms of misalignment” – in vitro demonstrations of the kinds of failures that might pose existential threats – as a new and importan
Explore this link on the map →saved by
related reading
- Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forumalignmentforum.org
- Advice for making robust-to-training model organismsblog.redwoodresearch.org
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- 39 - Evan Hubinger on Model Organisms of Misalignment | AXRP - the AI X-risk Research Podcastaxrp.net
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Your Model Organisms Might Be Fried — LessWronglesswrong.com
- Model Organisms for Emergent Misalignment — LessWronglesswrong.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Research Areas in Evaluation and Guarantees in Reinforcement Learning (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Chapter 4: Alignment Science - ARENAlearn.arena.education
- Alignment remains a hard, unsolved problem — AI Alignment Forumalignmentforum.org