The limits of AI safety via debate - LessWrong
The limits of AI safety via debate • I recently participated in the AGI safety fundamentals program and this is my cornerstone project. During our re…
x The limits of AI safety via debate — LessWrong Debate (AI safety technique) AI Frontpage 36 The limits of AI safety via debate by Marius Hobbhahn 10th May 2022 AI Alignment Forum 11 min read 8 36 Ω 17 The limits of AI safety via debate I recently participated in the AGI safety fundamentals program and this is my cornerstone project. During our readings of AI safety via debate ( blog , paper ) we had an interesting discussion on its limits and conditions under which it would fail. I spent only around 5 hours writing this post and it should thus mostly be seen as food for thought rather than r
Explore this link on the map →related reading
- [1805.00899] AI safety via debatear5iv.labs.arxiv.org
- [1805.00899] AI safety via debatearxiv.org
- An alignment safety case sketch based on debate — LessWronglesswrong.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- Anthropic Fall 2023 Debate Progress Update — AI Alignment Forumalignmentforum.org
- Optimal play in human-judged Debate usually won't answer your question — AI Alignment Forumalignmentforum.org
- Where I agree and disagree with Eliezer — LessWronglesswrong.com
- Avoiding Obfuscation with Prover-Estimator Debatearxiv.org
- AI Safety Seems Hard to Measurecold-takes.com
- AI Debaters are More Persuasive when Arguing in Alignment with Their Own Beliefsarxiv.org
- AI Safety for Fleshy Humans: a whirlwind touraisafety.dance
- Neel Nanda on Mechanistic Interpretability: Progress, Limits, and Paths to Safer AI — EA Forumforum.effectivealtruism.org