The generalization phase diagram — LessWrong
Introduction This is part I of 3.5 planned posts on the “tempered posterior” (one of which will be “SLT in a nutshell”). It is a kind of moral sequel…
x The generalization phase diagram — LessWrong Interpretability (ML & AI) Singular Learning Theory AI Frontpage 28 The generalization phase diagram by Dmitry Vaintrob 26th Jan 2025 20 min read 2 28 Introduction This is part I of 3.5 planned posts on the “tempered posterior” (one of which will be “SLT in a nutshell”). It is a kind of moral sequel to “ Dmitry’s Koan ” and “ Logits, log-odds, and loss for parallel circuits ”, and is related to the post on grammars . It can be read independently of any of these. In my view, one of the most valuable contributions of singular learning theory (SLT) s
Explore this link on the map →saved by
related reading
- arxiv.org/pdf/1805.08522arxiv.org
- AlgZoo: uninterpreted models with fewer than 1,500 parameters — LessWronglesswrong.com
- Machine Learning, Kolmogorov Complexity, and Squishy Bunniestheorangeduck.com
- On neural scaling and the quanta hypothesisericjmichaud.com
- DSLT 0. Distilling Singular Learning Theory — LessWronglesswrong.com
- Clare Lyle | What's grokking good for?clarelyle.com
- Deep learning as program synthesis — LessWronglesswrong.com
- Growth and Form in a Toy Model of Superposition — LessWronglesswrong.com
- Neural networks generalize because of this one weird trick — LessWronglesswrong.com
- Ambiguous out-of-distribution generalization on an algorithmic task — LessWronglesswrong.com
- Investigating the learning coefficient of modular addition: hackathon project — LessWronglesswrong.com
- Zipfian grokking | Jasper Gilleyjagilley.github.io