✳flâneur — a map of the web's best reading
New ARENA material: 8 exercise sets on alignment science & interpretability — LessWrong
lesswrong.com · 2,249 words · saved by 1 readers
TLDR This is a post announcing a lot of new ARENA material I've been working on for a while, which is now available for study here (currently on the…
x New ARENA material: 8 exercise sets on alignment science & interpretability — LessWrong Exercises / Problem-Sets AI Frontpage 2026 Top Fifty: 9 % 105 New ARENA material: 8 exercise sets on alignment science & interpretability by CallumMcDougall 27th Feb 2026 8 min read 1 105 TLDR This is a post announcing a lot of new ARENA material I've been working on for a while, which is now available for study here (currently on the alignment-science branch, but planned to be merged into main this Sunday). There's a set of exercises (each one contains about 1-2 days of material) on the following topics:
Explore this link on the map →related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- LessWronglesswrong.com
- ARENA - AI Safety Curriculumlearn.arena.education
- Chapter 4: Alignment Science - ARENAlearn.arena.education
- Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forumalignmentforum.org
- Teaching Claude why \ Anthropicanthropic.com
- Chapter 4: Alignment Science - ARENAlearn.arena.education
- How Can Interpretability Researchers Help AGI Go Well? — AI Alignment Forumalignmentforum.org
- 39 - Evan Hubinger on Model Organisms of Misalignment | AXRP - the AI X-risk Research Podcastaxrp.net
- GitHub - callummcdougall/ARENA_3.0 · GitHubgithub.com
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forumalignmentforum.org