How training-gamers might function (and win)
blog.redwoodresearch.org · 4,462 words · saved by 1 readers
A model of the relationship between higher level goals, explicit reasoning, and learned heuristics in capable agents.
In this post I present a model of the relationship between higher level goals, explicit reasoning, and learned heuristics in capable agents. This model suggests that given sufficiently rich training environments (and sufficient reasoning ability), models which terminally value on-episode reward-proxies are disadvantaged relative to training-gamers. A key point is that training gamers can still contain large quantities of learned heuristics (context-specific drives). By viewing these drives as instrumental and having good instincts for when to trust them, a training-gamer can capture the…
saved by
related reading
- When does training a model change its goals?blog.redwoodresearch.org
- Reward is not the optimization target — LessWronglesswrong.com
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- The behavioral selection model for predicting AI motivations — LessWronglesswrong.com
- A Toy Environment For Exploring Reasoning About Reward — LessWronglesswrong.com
- Models Don't "Get Reward" — LessWronglesswrong.com
- Specification gaming: the flip side of AI ingenuity — Google DeepMinddeepmind.google
- Reward Is Not the Optimization Targetturntrout.com
- Scheming AIs Will AIs fake alignment during training in order to get power?arxiv.org
- Thinking about reasoning models made me less worried about scheming — LessWronglesswrong.com
- Thinking about reasoning models made me less worried about scheming — LessWronglesswrong.com
- New report: “Scheming AIs: Will AIs fake alignment during training in order to get power?”joecarlsmith.com