SLT for AI Safety – Jesse Hoogland
jessehoogland.com · 1,023 words · saved by 1 readers
Jesse Hoogland's personal website
SLT for AI Safety – Jesse Hoogland SLT for AI Safety This sequence draws from a position paper co-written with Simon Pepin Lehalleur, Jesse Hoogland, Matthew Farrugia-Roberts, Susan Wei, Alexander Gietelink Oldenziel, Stan van Wingerden, George Wang, Zach Furman, Liam Carroll, Daniel Murfet. Thank you to Stan, Dan, and Simon for providing feedback on this post. **Alignment **\(\subseteq\) Capabilities. As of 2025, there is essentially no difference between the methods we use to align models and the methods we use to make models more capable. Everything is based on deep learning, and the main d
saved by
related reading
- SLT for AI Safety — LessWronglesswrong.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Teaching Claude why \ Anthropicanthropic.com
- Subliminal Learning: Language Models Transmit Behavioral Traits via Hidden Signals in Dataalignment.anthropic.com
- Timaeus | Learn about SLTtimaeus.co
- Iliad Intensive Curriculumiliad-team.github.io
- Iliad Intensive Curriculumiliad-intensive.org
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- The Iliad Intensive Course Materials — LessWronglesswrong.com
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- Safe AI with Singular Learning Theory ..hyper-exponential.com
- Why AI alignment could be hard with modern deep learningcold-takes.com