Models Don't "Get Reward" - LessWrong
In terms of content, this has a lot of overlap with Reward is not the optimization target. I'm basically rewriting a part of that post in language I personally find clearer, emphasising what I think…
x Models Don't "Get Reward" — LessWrong Best of LessWrong 2022 Reinforcement learning Distillation & Pedagogy Goal-Directedness AI Curated 349 Models Don't "Get Reward" by Sam Ringer 30th Dec 2022 AI Alignment Forum 6 min read 64 349 Ω 84 In terms of content, this has a lot of overlap with Reward is not the optimization target . I'm basically rewriting a part of that post in language I personally find clearer, emphasising what I think is the core insight. When thinking about deception and RLHF training, a simplified threat model is something like this: A model takes some actions. If a human ap
Explore this link on the map →saved by
related reading
- Reward is not the optimization target — LessWronglesswrong.com
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Reward is not the optimization target — AI Alignment Forumalignmentforum.org
- The behavioral selection model for predicting AI motivations — LessWronglesswrong.com
- [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Modelarxiv.org
- Reward Is Not Enough — LessWronglesswrong.com
- Reward Is Not the Optimization Targetturntrout.com
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- Evolution as Backstop for Reinforcement Learning · Gwern.netgwern.net
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- (Some) Natural Emergent Misalignment from Reward Hacking in Non-Production RL — LessWronglesswrong.com
- Reward hacking is becoming more sophisticated and deliberate in frontier LLMs — LessWronglesswrong.com