Self-prediction acts as an emergent regularizer — AI Alignment Forum
In our recent work with Professor Michael Graziano (arXiv, thread), we show that adding an auxiliary self-modeling objective to supervised learning tasks yields simpler, more regularized, and more parameter-efficient models. Across three classification tasks and two modalities, self-modeling consistently reduced complexity (lower RLCT, narrower weight distribution). This restructuring effect may help explain the putative benefits of self-models in both ML and biological systems. Agents who self-model may be reparameterized to better predict themselves, predict others, and be predicted by others. Accordingly, we believe that further exploring the potential effects of self-modeling on cooperation emerges as a promising neglected approach to alignment. This approach may also exhibit a 'negative alignment tax' to the degree that it may end up enhancing alignment and rendering systems more globally effective. In this post, we discuss some of the core findings and implications of our recent
x Self-prediction acts as an emergent regularizer — AI Alignment Forum Consciousness Research Agendas AI Frontpage 20 Self-prediction acts as an emergent regularizer by Cameron Berg , Kvee , Mike Vaiana , Diogo de Lucena , florin_pop , Trent Hodgeson 23rd Oct 2024 5 min read 9 20 TL;DR: In our recent work with Professor Michael Graziano ( arXiv , thread ), we show that adding an auxiliary self-modeling objective to supervised learning tasks yields simpler, more regularized, and more parameter-efficient models. Across three classification tasks and two modalities, self-modeling consistently red
Explore this link on the map →related reading
- Self-prediction acts as an emergent regularizer — LessWronglesswrong.com
- Unexpected Benefits of Self-Modeling in Neural Systemsarxiv.org
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Simulators — LessWronglesswrong.com
- Self-Fulfilling Misalignment Data Might Be Poisoning Our AI Modelsturntrout.com
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- Position: It's Time to Optimize for Self-Consistencytime-for-consistency.github.io
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- DL towards the unaligned Recursive Self-Optimization attractor — LessWronglesswrong.com
- A case for LLMs as Self-predictors — LessWronglesswrong.com