Unexpected Benefits of Self-Modeling in Neural Systems
Self-models have been a topic of great interest for decades in studies of human cognition and more recently in machine learning. Yet what benefits do self-models confer? Here we show that when artificial networks learn to predict their internal states as an auxiliary task, they change in a fundamental way. To better perform the self-model task, the network learns to make itself simpler, more regularized, more parameter-efficient, and therefore more amenable to being predictively modeled. To test the hypothesis of self-regularizing through self-modeling, we used a range of network architectures performing three classification tasks across two modalities. In all cases, adding self- modeling caused a significant reduction in network complexity. The reduction was observed in two ways. First, the distribution of weights was narrower when self-modeling was present. Second, a measure of network complexity, the real log canonical threshold (RLCT), was smaller when self-modeling was present. No
Unexpected Benefits of Self-Modeling in Neural Systems Vickram N. Premakumar Michael Vaiana Florin Pop Judd Rosenblatt Diogo Schwerz de Lucena AE Studio, Venice, CA Kirsten Ziman Princeton Neuroscience Institute, Princeton University, Princeton NJ Michael S. A. Graziano graziano@princeton.edu Princeton Neuroscience Institute, Princeton University, Princeton NJ Department of Psychology, Princeton University, Princeton, NJ (July 14, 2024) Abstract Self-models have been a topic of great interest for decades in studies of human cognition and more recently in machine learning. Yet what benefits do
Explore this link on the map →related reading
- Self-prediction acts as an emergent regularizer — LessWronglesswrong.com
- Self-prediction acts as an emergent regularizer — AI Alignment Forumalignmentforum.org
- Toy Models of Superpositiontransformer-circuits.pub
- Aman's AI Journal • Primers • Ilya Sutskever's Top 30aman.ai
- On neural scaling and the quanta hypothesisericjmichaud.com
- pdfopenreview.net
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- NL.pdfabehrouz.github.io
- AlgZoo: uninterpreted models with fewer than 1,500 parameters — LessWronglesswrong.com
- DL towards the unaligned Recursive Self-Optimization attractor — LessWronglesswrong.com
- Position: It's Time to Optimize for Self-Consistencytime-for-consistency.github.io
- Researchbactra.org