Alignment likely generalizes further than capabilities.
Recently, I was reading this paper which demonstrates how to do online RLHF for alignment of LLMs and a sentence stuck out to me: We conjecture that this is because the reward model (discriminator) usually generalizes better than the policy (generator) This is an offhand remark but it strikes at...
Recently, I was reading this paper which demonstrates how to do online RLHF for alignment of LLMs and a sentence stuck out to me: We conjecture that this is because the reward model (discriminator) usually generalizes better than the policy (generator) This is an offhand remark but it strikes at the core of the LessWrong doom story, as propounded here for which the idea of ‘capabilities generalizing further than alignment’ is central. This is a key point because it is necessary for the self-improvmeent of AI systems, even if initially ‘aligned’ using current methods, to spell doom because even
Explore this link on the map →related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Will Capabilities Generalise More? — AI Alignment Forumalignmentforum.org
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Why I’m optimistic about our alignment approachaligned.substack.com
- A Toy Environment For Exploring Reasoning About Reward — LessWronglesswrong.com
- Thomas Larsen's Shortform — LessWronglesswrong.com
- Research Areas in Evaluation and Guarantees in Reinforcement Learning (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- Models Don't "Get Reward" — LessWronglesswrong.com
- Why I’m optimistic about our alignment approachaligned.substack.com
- Alignment remains a hard, unsolved problem — AI Alignment Forumalignmentforum.org
- Research Areas in Methods for Post-training and Elicitation (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org