Anthropic Fall 2023 Debate Progress Update — AI Alignment Forum
This is a research update on some work that I’ve been doing on Scalable Oversight at Anthropic, based on the original AI safety via debate proposal a…
x Anthropic Fall 2023 Debate Progress Update — AI Alignment Forum Debate (AI safety technique) World Modeling Frontpage 32 Anthropic Fall 2023 Debate Progress Update by Ansh Radhakrishnan 28th Nov 2023 15 min read 9 32 This is a research update on some work that I’ve been doing on Scalable Oversight at Anthropic, based on the original AI safety via debate proposal and a more recent agenda developed at NYU and Anthropic. The core doc was written several months ago, so some of it is likely outdated, but it seemed worth sharing in its current form. I’d like to thank Tamera Lanham, Sam Bowman, Kam
related reading
- [1805.00899] AI safety via debatear5iv.labs.arxiv.org
- The limits of AI safety via debate — LessWronglesswrong.com
- AI Debaters are More Persuasive when Arguing in Alignment with Their Own Beliefsarxiv.org
- [1805.00899] AI safety via debatearxiv.org
- [2506.13609] Avoiding Obfuscation with Prover-Estimator Debatearxiv.org
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- Optimal play in human-judged Debate usually won't answer your question — AI Alignment Forumalignmentforum.org
- An alignment safety case sketch based on debate — LessWronglesswrong.com
- Teaching Claude Whyalignment.anthropic.com
- Deep Deceptiveness — LessWronglesswrong.com
- Agentic Misalignment in Summer 2026alignment.anthropic.com