Investigating the learning coefficient of modular addition: hackathon project — LessWrong
As our project at the Melbourne hackathon on Singular Learning Theory and alignment (Oct. 7-8), we did some experiments to estimate the learning coef…
x Investigating the learning coefficient of modular addition: hackathon project — LessWrong Singular Learning Theory Interpretability (ML & AI) AI Frontpage 97 Investigating the learning coefficient of modular addition: hackathon project by Nina Panickssery , Dmitry Vaintrob 17th Oct 2023 AI Alignment Forum 15 min read 5 97 Ω 42 As our project at the Melbourne hackathon on Singular Learning Theory and alignment (Oct. 7-8), we did some experiments to estimate the learning coefficient of the single-layer modular addition task at a basin, an invariant that measures the information complexity (rea
Explore this link on the map →saved by
related reading
- DSLT 0. Distilling Singular Learning Theory — LessWronglesswrong.com
- On neural scaling and the quanta hypothesisericjmichaud.com
- AlgZoo: uninterpreted models with fewer than 1,500 parameters — LessWronglesswrong.com
- Neural networks generalize because of this one weird trick — LessWronglesswrong.com
- Neural networks generalize because of this one weird trick — AI Alignment Forumalignmentforum.org
- LoRA Without Regret - Thinking Machines Labthinkingmachines.ai
- [2602.16849] On the Mechanism and Dynamics of Modular Addition: Fourier Features, Lottery Ticket, and Grokkingarxiv.org
- Timaeus | Learn about SLTtimaeus.co
- The Scaling Hypothesis · Gwern.netgwern.net
- [2604.21691] There Will Be a Scientific Theory of Deep Learningarxiv.org
- Clare Lyle | What's grokking good for?clarelyle.com
- Do Machine Learning Models Memorize or Generalize?pair.withgoogle.com