ARENA - AI Safety Curriculum
This is where the ARENA course content is hosted. For more information about the ARENA program, including upcoming cohorts and how to apply, visit arena.education.
ARENA - AI Safety Curriculum 0 Fundamentals Build your foundation in deep learning, from prerequisites through CNNs, optimization, backpropagation, and generative models. 6 sections 1 Interpretability Dive deep into language model interpretability, from linear probes and SAEs to circuit analysis and toy models. 13 sections 2 RL Take a whirlwind tour through RL, starting from tabular learning and Atari, and ending with some of the cutting-edge techniques used in current LLM post-training. 5 sections 3 Evals Learn to build and run evaluations for large language models, including dataset generati
Explore this link on the map →saved by
related reading
- Transformer Circuits Threadtransformer-circuits.pub
- The Iliad Intensive Course Materials — LessWronglesswrong.com
- aman.ai • the art of artificial intelligenceaman.ai
- [2401.05566] Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Trainingarxiv.org
- Neuronpedianeuronpedia.org
- GenAI Handbookgenai-handbook.github.io
- GitHub - jacobhilton/deep_learning_curriculum: Language model alignment-focused deep learning curriculum · GitHubgithub.com
- GitHub - callummcdougall/ARENA_3.0 · GitHubgithub.com
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- AI in 2025: gestalt — LessWronglesswrong.com
- AI Safety | Arkosevictoriabrook.github.io
- ARENA 6.0 Impact Report — LessWronglesswrong.com