AI Alignment
This curriculum was developed with Richard Ngo (Open AI) and was contributed to by multiple experts in the field of alignment. We expect each section takes 2-4 hours to engage with all the materials (16-32 hours overall). Additionally, there are exercises to help you think through topics yourself and make progress in your learning about AI alignment. We organise content into weeks to help you pace your engagement with the curriculum. By the end of this course, you should be able to understand a range of agenda in AI Alignment and make informed decisions about your next steps to engage with the field.
How do we ensure advanced AI systems act in line with human intentions? Apply now Curriculum AI and the years ahead Unit 1 View unit→ What is AI alignment? Unit 2 View unit→ Reinforcement learning from human (or AI) feedback Unit 3 View unit→ Scalable oversight Unit 4 View unit→ Robustness unlearning and control Unit 5 View unit→ Mechanistic interpretability Unit 6 View unit→ Technical governance approaches Unit 7 View unit→ Contributing to AI safety Unit 8 View unit→ Rapidly testing your project Unit 9 View unit→ Developing your project Unit 10 View unit→…
saved by
related reading
- AGI safety career advice — EA Forumforum.effectivealtruism.org
- Frontier AI Governance Course | BlueDot Impactcourse.aisafetyfundamentals.com
- AI Alignment | BlueDot Impactcourse.aisafetyfundamentals.com
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- Fall 2026boazbk.github.io
- Study Guide — LessWronglesswrong.com
- The Iliad Intensive Course Materials — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- A woefully incomplete guide to technical upskillingjason.ml
- GitHub - jacobhilton/deep_learning_curriculum: Language model alignment-focused deep learning curriculumgithub.com
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com