An overview of 11 proposals for building safe advanced AI — AI Alignment Forum
This is the blog post version of the paper by the same name. Special thanks to Kate Woolverton, Paul Christiano, Rohin Shah, Alex Turner, William Sau…
x An overview of 11 proposals for building safe advanced AI — AI Alignment Forum Best of LessWrong 2020 Research Agendas AI Success Models Scalable Oversight AI Risk Debate (AI safety technique) Inner Alignment Iterated Amplification Myopia Outer Alignment AI Curated 72 An overview of 11 proposals for building safe advanced AI by evhub 29th May 2020 46 min read 37 72 This is the blog post version of the paper by the same name . Special thanks to Kate Woolverton, Paul Christiano, Rohin Shah, Alex Turner, William Saunders, Beth Barnes, Abram Demski, Scott Garrabrant, Sam Eisenstat, and Tsvi Bens
Explore this link on the map →related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- Research Areas in Evaluation and Guarantees in Reinforcement Learning (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- Relaxed adversarial training for inner alignment — AI Alignment Forumalignmentforum.org
- Research Areas in Methods for Post-training and Elicitation (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- An AI alignment research agenda based on asymmetric debate and monitoring. — LessWronglesswrong.com
- Why I’m optimistic about our alignment approachaligned.substack.com
- Alignment remains a hard, unsolved problem — AI Alignment Forumalignmentforum.org
- Externalized reasoning oversight: a research direction for language model alignment — AI Alignment Forumalignmentforum.org