Models Don't "Get Reward" - LessWrong
In terms of content, this has a lot of overlap with Reward is not the optimization target. I'm basically rewriting a part of that post in language I personally find clearer, emphasising what I think…
x Models Don't "Get Reward" — LessWrong Best of LessWrong 2022 Reinforcement learning Distillation & Pedagogy Goal-Directedness AI Curated 349 Models Don't "Get Reward" by Sam Ringer 30th Dec 2022 AI Alignment Forum 6 min read 64 349 Ω 84 In terms of content, this has a lot of overlap with Reward is not the optimization target . I'm basically rewriting a part of that post in language I personally find clearer, emphasising what I think is the core insight. When thinking about deception and RLHF training, a simplified threat model is something like this: A model takes some actions. If a human ap
saved by
related reading
- Reward is not the optimization target — LessWronglesswrong.com
- Reward Is Not the Optimization Targetturntrout.com
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Reward is not the optimization target — AI Alignment Forumalignmentforum.org
- The behavioral selection model for predicting AI motivations — LessWronglesswrong.com
- [2607.18966] Measuring Reward-Seeking via Contrastive Belief Updatesarxiv.org
- [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Modelarxiv.org
- Reward Is Not Enough — LessWronglesswrong.com
- Measuring Reward-Seeking by Instilling Contrastive Beliefsalignment.openai.com
- Evolution as Backstop for Reinforcement Learning · Gwern.netgwern.net
- (Some) Natural Emergent Misalignment from Reward Hacking in Non-Production RL — LessWronglesswrong.com
- RLHF | John Lambertjohnwlambert.github.io