Anthropic Fall 2023 Debate Progress Update — AI Alignment Forum
This is a research update on some work that I’ve been doing on Scalable Oversight at Anthropic, based on the original AI safety via debate proposal a…
x Anthropic Fall 2023 Debate Progress Update — AI Alignment Forum Debate (AI safety technique) World Modeling Frontpage 32 Anthropic Fall 2023 Debate Progress Update by Ansh Radhakrishnan 28th Nov 2023 15 min read 9 32 This is a research update on some work that I’ve been doing on Scalable Oversight at Anthropic, based on the original AI safety via debate proposal and a more recent agenda developed at NYU and Anthropic. The core doc was written several months ago, so some of it is likely outdated, but it seemed worth sharing in its current form. I’d like to thank Tamera Lanham, Sam Bowman, Kam
Explore this link on the map →related reading
- [1805.00899] AI safety via debatear5iv.labs.arxiv.org
- AI Debaters are More Persuasive when Arguing in Alignment with Their Own Beliefsarxiv.org
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- Optimal play in human-judged Debate usually won't answer your question — AI Alignment Forumalignmentforum.org
- An alignment safety case sketch based on debate — LessWronglesswrong.com
- The limits of AI safety via debate — LessWronglesswrong.com
- [1805.00899] AI safety via debatearxiv.org
- Teaching Claude why \ Anthropicanthropic.com
- Externalized reasoning oversight: a research direction for language model alignment — AI Alignment Forumalignmentforum.org
- Sandbagging with misaligned action - Chain-of-Thought Transcript - Anti-Schemingantischeming.ai
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com