Optimal play in human-judged Debate usually won't answer your question - AI Alignment Forum
Note: I fully expect some readers to find the core of this post almost trivially obvious. If you’re such a reader, please read as “I think [obvious thing] is important”, rather than “I’ve discovered…
x Optimal play in human-judged Debate usually won't answer your question — AI Alignment Forum Debate (AI safety technique) AI Rationality Frontpage 20 Optimal play in human-judged Debate usually won't answer your question by Joe Collman 27th Jan 2021 15 min read 12 20 Epistemic status: highly confident (99%+) this is an issue for optimal play with human consequentialist judges. Thoughts on practical implications are more speculative, and involve much hand-waving (70% sure I’m not overlooking a trivial fix, and that this can’t be safely ignored). Note: I fully expect some readers to find the co
Explore this link on the map →related reading
- [1805.00899] AI safety via debatear5iv.labs.arxiv.org
- Anthropic Fall 2023 Debate Progress Update — AI Alignment Forumalignmentforum.org
- [1805.00899] AI safety via debatearxiv.org
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- The limits of AI safety via debate — LessWronglesswrong.com
- AI Debaters are More Persuasive when Arguing in Alignment with Their Own Beliefsarxiv.org
- An alignment safety case sketch based on debate — LessWronglesswrong.com
- Optimality is the tiger, and agents are its teeth — LessWronglesswrong.com
- Everything I have to say about debate - by Eli Qianeliqian.substack.com
- The Best of LessWrong — LessWronglesswrong.com
- AI for Decision Adviceforethought.org
- The Best of LessWrong — LessWronglesswrong.com