flâneur — a map of the web's best reading

Models Don't "Get Reward" - LessWrong

lesswrong.com · 12,576 words · saved by 2 readers

In terms of content, this has a lot of overlap with Reward is not the optimization target. I'm basically rewriting a part of that post in language I personally find clearer, emphasising what I think…

x Models Don't "Get Reward" — LessWrong Best of LessWrong 2022 Reinforcement learning Distillation & Pedagogy Goal-Directedness AI Curated 349 Models Don't "Get Reward" by Sam Ringer 30th Dec 2022 AI Alignment Forum 6 min read 64 349 Ω 84 In terms of content, this has a lot of overlap with Reward is not the optimization target . I'm basically rewriting a part of that post in language I personally find clearer, emphasising what I think is the core insight. When thinking about deception and RLHF training, a simplified threat model is something like this: A model takes some actions. If a human ap

Explore this link on the map →

saved by

related reading