flâneur — a map of the web's best reading

SLT for AI Safety — LessWrong

lesswrong.com · 1,092 words · saved by 1 readers

> This sequence draws from a position paper co-written with Simon Pepin Lehalleur, Jesse Hoogland, Matthew Farrugia-Roberts, Susan Wei, Alexander Gie…

x SLT for AI Safety — LessWrong SLT for AI Safety Inner Alignment Interpretability (ML & AI) Singular Learning Theory AI Frontpage 78 SLT for AI Safety by Jesse Hoogland 1st Jul 2025 AI Alignment Forum 4 min read 0 78 Ω 21 This sequence draws from a position paper co-written with Simon Pepin Lehalleur, Jesse Hoogland, Matthew Farrugia-Roberts, Susan Wei, Alexander Gietelink Oldenziel, Stan van Wingerden, George Wang, Zach Furman, Liam Carroll, Daniel Murfet. Thank you to Stan, Dan, and Simon for providing feedback on this post. Alignment ⊆ Capabilities. As of 2025, there is essentially no diff

Explore this link on the map →

saved by

related reading