SLT for AI Safety — LessWrong
> This sequence draws from a position paper co-written with Simon Pepin Lehalleur, Jesse Hoogland, Matthew Farrugia-Roberts, Susan Wei, Alexander Gie…
x SLT for AI Safety — LessWrong SLT for AI Safety Inner Alignment Interpretability (ML & AI) Singular Learning Theory AI Frontpage 78 SLT for AI Safety by Jesse Hoogland 1st Jul 2025 AI Alignment Forum 4 min read 0 78 Ω 21 This sequence draws from a position paper co-written with Simon Pepin Lehalleur, Jesse Hoogland, Matthew Farrugia-Roberts, Susan Wei, Alexander Gietelink Oldenziel, Stan van Wingerden, George Wang, Zach Furman, Liam Carroll, Daniel Murfet. Thank you to Stan, Dan, and Simon for providing feedback on this post. Alignment ⊆ Capabilities. As of 2025, there is essentially no diff
Explore this link on the map →saved by
related reading
- Timaeus | Learn about SLTtimaeus.co
- SLT for AI Safety – Jesse Hooglandjessehoogland.com
- Maybe I was too harsh on deep learning theory (three days ago) — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- The Iliad Intensive Course Materials — LessWronglesswrong.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Why AI alignment could be hard with modern deep learningcold-takes.com
- Intentionally Designing the Future of AIgoodfire.ai
- Subliminal Learning: Language Models Transmit Behavioral Traits via Hidden Signals in Dataalignment.anthropic.com
- Teaching Claude why \ Anthropicanthropic.com
- ARENA - AI Safety Curriculumlearn.arena.education