Distilling Singular Learning Theory - LessWrong
This sequence distills Sumio Watanabe's Singular Learning Theory (SLT) by explaining the essence of its main theorem - Watanabe's Free Energy Formula for Singular Models - and illustrating its implications with intuition-building examples. I show why neural networks are singular models, and demonstrate how SLT provides a framework for understanding phases and phase transitions in neural networks.
x Distilling Singular Learning Theory — LessWrong Distilling Singular Learning Theory Jun 16, 2023 by Liam Carroll This sequence distills Sumio Watanabe's Singular Learning Theory (SLT) by explaining the essence of its main theorem - Watanabe's Free Energy Formula for Singular Models - and illustrating its implications with intuition-building examples. I show why neural networks are singular models, and demonstrate how SLT provides a framework for understanding phases and phase transitions in neural networks. 97 DSLT 0. Distilling Singular Learning Theory Ω Liam Carroll 3y Ω 8 62 DSLT 1. The R
Explore this link on the map →related reading
- DSLT 0. Distilling Singular Learning Theory — LessWronglesswrong.com
- DSLT 2. Why Neural Networks obey Occam's Razor — LessWronglesswrong.com
- DSLT 1. The RLCT Measures the Effective Dimension of Neural Networks — LessWronglesswrong.com
- DSLT 3. Neural Networks are Singular — LessWronglesswrong.com
- DSLT 2. Why Neural Networks obey Occam's Razor — LessWronglesswrong.com
- Timaeus | Learn about SLTtimaeus.co
- Neural networks generalize because of this one weird trick — LessWronglesswrong.com
- Neural networks generalize because of this one weird trick — AI Alignment Forumalignmentforum.org
- Maybe I was too harsh on deep learning theory (three days ago) — LessWronglesswrong.com
- On neural scaling and the quanta hypothesisericjmichaud.com
- [2604.21691] There Will Be a Scientific Theory of Deep Learningarxiv.org
- Toy Models of Superpositiontransformer-circuits.pub