An AI alignment research agenda based on asymmetric debate and monitoring. — LessWrong
TL;DR: Under cooperative conditions and an assumption that we have to build ASI in relatively short timelines, we might solve alignment by building s…
x An AI alignment research agenda based on asymmetric debate and monitoring. — LessWrong Debate (AI safety technique) Research Agendas AI Community World Optimization Frontpage 4 An AI alignment research agenda based on asymmetric debate and monitoring. by emanuelr 10th Apr 2026 22 min read 0 4 Epistemic Status: Personal research agenda exploring alignment approaches under assumptions of human coordination and bounded timelines. Not a comprehensive survey reflects my interests and current thinking. TL;DR: Under cooperative conditions and an assumption that we have to build ASI in relatively sh
Explore this link on the map →related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- Alignment remains a hard, unsolved problem — AI Alignment Forumalignmentforum.org
- Research Areas in Evaluation and Guarantees in Reinforcement Learning (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- Why I’m optimistic about our alignment approachaligned.substack.com
- Externalized reasoning oversight: a research direction for language model alignment — AI Alignment Forumalignmentforum.org
- (My understanding of) What Everyone in Technical Alignment is Doing and Why — LessWronglesswrong.com
- Why I’m optimistic about our alignment approachaligned.substack.com