Self-prediction acts as an emergent regularizer — LessWrong
In our recent work with Professor Michael Graziano (arXiv, thread), we show that adding an auxiliary self-modeling objective to supervised learning tasks yields simpler, more regularized, and more parameter-efficient models. Across three classification tasks and two modalities, self-modeling consistently reduced complexity (lower RLCT, narrower weight distribution). This restructuring effect may help explain the putative benefits of self-models in both ML and biological systems. Agents who self-model may be reparameterized to better predict themselves, predict others, and be predicted by others. Accordingly, we believe that further exploring the potential effects of self-modeling on cooperation emerges as a promising neglected approach to alignment. This approach may also exhibit a 'negative alignment tax' to the degree that it may end up enhancing alignment and rendering systems more globally effective. In this post, we discuss some of the core findings and implications of our recent
x Self-prediction acts as an emergent regularizer — LessWrong Consciousness Research Agendas AI Frontpage 92 Self-prediction acts as an emergent regularizer by Cameron Berg , Kvee , Mike Vaiana , Diogo de Lucena , florin_pop , Trent Hodgeson 23rd Oct 2024 AI Alignment Forum 5 min read 9 92 Ω 20 TL;DR: In our recent work with Professor Michael Graziano ( arXiv , thread ), we show that adding an auxiliary self-modeling objective to supervised learning tasks yields simpler, more regularized, and more parameter-efficient models. Across three classification tasks and two modalities, self-modeling c
Explore this link on the map →related reading
- Self-prediction acts as an emergent regularizer — AI Alignment Forumalignmentforum.org
- Unexpected Benefits of Self-Modeling in Neural Systemsarxiv.org
- Aman's AI Journal • Primers • Ilya Sutskever's Top 30aman.ai
- Simulators — LessWronglesswrong.com
- Position: It's Time to Optimize for Self-Consistencytime-for-consistency.github.io
- pdfopenreview.net
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- Self-Fulfilling Misalignment Data Might Be Poisoning Our AI Modelsturntrout.com
- DL towards the unaligned Recursive Self-Optimization attractor — LessWronglesswrong.com
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- A case for LLMs as Self-predictors — LessWronglesswrong.com
- AlgZoo: uninterpreted models with fewer than 1,500 parameters — LessWronglesswrong.com