✳flâneur — a map of the web's best reading
Iterated Distillation and Amplification | by Ajeya Cotra | AI Alignment
ai-alignment.com · 1,945 words · saved by 1 readers
Guest post summarizing my approach to aligned RL.
Machine Learning Iterated Distillation and Amplification Ajeya Cotra 8 min read · Mar 4, 2018 -- 6 Listen Share This is a guest post summarizing Paul Christiano’s proposed scheme for training machine learning systems that can be robustly aligned to complex and fuzzy values, which I call Iterated Distillation and Amplification (IDA) here. IDA is notably similar to AlphaGoZero and expert iteration . The hope is that if we use IDA to train each learned component of an AI then the overall AI will remain aligned with the user’s interests while achieving state of the art performance at runtime — pro
Explore this link on the map →related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- A guide to Iterated Amplification & Debate — AI Alignment Forumalignmentforum.org
- Teaching Claude why \ Anthropicanthropic.com
- AI safety techniques leveraging distillation — LessWronglesswrong.com
- Supervising strong learners by amplifying weak experts — AI Alignment Forumalignmentforum.org
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- The Iliad Intensive Course Materials — LessWronglesswrong.com
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- AI 2027ai-2027.com
- [1810.08575] Supervising strong learners by amplifying weak expertsar5iv.labs.arxiv.org
- [AN #70]: Agents that help humans who are still learning about their own preferences — LessWronglesswrong.com
- Thomas Larsen's Shortform — LessWronglesswrong.com