DSLT 3. Neural Networks are Singular — LessWrong
TLDR; This is the third main post of Distilling Singular Learning Theory which is introduced in DSLT0. I explain that neural networks are singular models because of the symmetries in parameter space that produce the same function, and introduce a toy two layer ReLU neural network setup where these symmetries can be perfectly classified. I provide motivating examples of each kind of symmetry, with particular emphasis on the non-generic node-degeneracy and orientation-reversing symmetries that give rise to interesting phases to be studied in DSLT4. As we discussed in DSLT2, singular models have the capacity to generalise well because the effective dimension of a singular model, as measured by the RLCT, can be less than half the dimension of parameter space. With this in mind, it should be no surprise that neural networks are indeed singular models, but up until this point we have not exactly explained what feature they possess that makes them singular. In this post, we will explain that
x DSLT 3. Neural Networks are Singular — LessWrong window.__lwSsrGql.inject("query PostsPageWrapper($documentId: String, $sequenceId: String) {\n post(input: {selector: {documentId: $documentId}}, allowNull: true) {\n result {\n ...PostsWithNavigation\n }\n }\n}\n\nfragment PostsMinimumInfo on Post {\n _id\n slug\n title\n draft\n shortform\n hideCommentKarma\n af\n userId\n coauthorUserIds\n rejected\n collabEditorDialogue\n}\n\nfragment PostsBase on Post {\n ...PostsMinimumInfo\n url\n postedAt\n sticky\n metaSticky\n stickyPriority\n status\n frontpageDate\n meta\n deletedDraft\n postCatego
Explore this link on the map →related reading
- DSLT 0. Distilling Singular Learning Theory — LessWronglesswrong.com
- Neural networks generalize because of this one weird trick — AI Alignment Forumalignmentforum.org
- Neural networks generalize because of this one weird trick — LessWronglesswrong.com
- Distilling Singular Learning Theory — LessWronglesswrong.com
- DSLT 1. The RLCT Measures the Effective Dimension of Neural Networks — LessWronglesswrong.com
- DSLT 2. Why Neural Networks obey Occam's Razor — LessWronglesswrong.com
- Toy Models of Superpositiontransformer-circuits.pub
- DSLT 2. Why Neural Networks obey Occam's Razor — LessWronglesswrong.com
- [2604.21691] There Will Be a Scientific Theory of Deep Learningarxiv.org
- Timaeus | Learn about SLTtimaeus.co
- Maybe I was too harsh on deep learning theory (three days ago) — LessWronglesswrong.com
- Neural Networks, Manifolds, and Topology -- colah's blogcolah.github.io