Alignment likely generalizes further than capabilities.
Recently, I was reading this paper which demonstrates how to do online RLHF for alignment of LLMs and a sentence stuck out to me: We conjecture that this is because the reward model (discriminator) usually generalizes better than the policy (generator) This is an offhand remark but it strikes at...
Recently, I was reading this paper which demonstrates how to do online RLHF for alignment of LLMs and a sentence stuck out to me: We conjecture that this is because the reward model (discriminator) usually generalizes better than the policy (generator) This is an offhand remark but it strikes at the core of the LessWrong doom story, as propounded here for which the idea of ‘capabilities generalizing further than alignment’ is central. This is a key point because it is necessary for the self-improvmeent of AI systems, even if initially ‘aligned’ using current methods, to spell doom because even
related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Alignment is not solved but it increasingly looks solvablealigned.substack.com
- Will Capabilities Generalise More? — AI Alignment Forumalignmentforum.org
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Reinforcement learning towards broadly and persistently beneficial modelsalignment.openai.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- RLHF | John Lambertjohnwlambert.github.io
- Why I’m optimistic about our alignment approachaligned.substack.com
- Narrow Misalignment is Hard, Emergent Misalignment is Easy — LessWronglesswrong.com
- A Toy Environment For Exploring Reasoning About Reward — LessWronglesswrong.com
- Thomas Larsen's Shortform — LessWronglesswrong.com
- Alignment Is Proven To Be Solvable - by SE Gygesverysane.ai