ARENA - AI Safety Curriculum
This is where the ARENA course content is hosted. For more information about the ARENA program, including upcoming cohorts and how to apply, visit arena.education.
ARENA - AI Safety Curriculum 0 Fundamentals Build your foundation in deep learning, from prerequisites through CNNs, optimization, backpropagation, and generative models. 6 sections 1 Interpretability Dive deep into language model interpretability, from linear probes and SAEs to circuit analysis and toy models. 13 sections 2 RL Take a whirlwind tour through RL, starting from tabular learning and Atari, and ending with some of the cutting-edge techniques used in current LLM post-training. 5 sections 3 Evals Learn to build and run evaluations for large language models, including dataset generati
saved by
related reading
- Transformer Circuits Threadtransformer-circuits.pub
- The Iliad Intensive Course Materials — LessWronglesswrong.com
- aman.ai • the art of artificial intelligenceaman.ai
- Papers and Projects - Josh Engelsjoshengels.com
- Fall 2026boazbk.github.io
- GenAI Handbookgenai-handbook.github.io
- GitHub - callummcdougall/ARENA_3.0 · GitHubgithub.com
- [2401.05566] Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Trainingarxiv.org
- Neuronpedianeuronpedia.org
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- CAMBRIA — Cambridge Boston Alignment Initiativecbai.ai
- GitHub - jacobhilton/deep_learning_curriculum: Language model alignment-focused deep learning curriculumgithub.com