Iterated Distillation and Amplification | by Ajeya Cotra | AI Alignment
ai-alignment.com · 1,945 words · saved by 1 readers
Guest post summarizing my approach to aligned RL.
Machine Learning Iterated Distillation and Amplification Ajeya Cotra 8 min read · Mar 4, 2018 -- 6 Listen Share This is a guest post summarizing Paul Christiano’s proposed scheme for training machine learning systems that can be robustly aligned to complex and fuzzy values, which I call Iterated Distillation and Amplification (IDA) here. IDA is notably similar to AlphaGoZero and expert iteration . The hope is that if we use IDA to train each learned component of an AI then the overall AI will remain aligned with the user’s interests while achieving state of the art performance at runtime — pro
related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- A guide to Iterated Amplification & Debate — AI Alignment Forumalignmentforum.org
- Teaching Claude why \ Anthropicanthropic.com
- AI safety techniques leveraging distillation — LessWronglesswrong.com
- Supervising strong learners by amplifying weak experts — AI Alignment Forumalignmentforum.org
- The Iliad Intensive Course Materials — LessWronglesswrong.com
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- Teaching Claude Whyalignment.anthropic.com
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- [1810.08575] Supervising strong learners by amplifying weak expertsar5iv.labs.arxiv.org
- The distillation double bind: Distilling misaligned models either transfers misalignment or it doesn'tblog.redwoodresearch.org
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com