“Behaviorist” RL reward functions lead to scheming — AI Alignment Forum
I will argue that a large class of reward functions, which I call “behaviorist”, and which includes almost every reward function in the RL and LLM literature, are all doomed to eventually lead to AI that will “scheme”—i.e., pretend to be docile and cooperative while secretly looking for opportunities to behave in egregiously bad ways such as world takeover (cf. “treacherous turn”). I’ll mostly focus on “brain-like AGI” (as defined just below), but I think the argument applies equally well to future LLMs, if their competence comes overwhelmingly from RL rather than from pretraining.[1] The issue is basically that “negative reward for lying and stealing” looks the same as “negative reward for getting caught lying and stealing”. I’ll argue that the AI will wind up with the latter motivation. The reward function will miss sufficiently sneaky misaligned behavior, and so the AI will come to feel like that kind of behavior is good, and this tendency will generalize in a very bad way. What ver
x “Behaviorist” RL reward functions lead to scheming — AI Alignment Forum Deceptive Alignment AI Frontpage 25 “Behaviorist” RL reward functions lead to scheming by Steven Byrnes 23rd Jul 2025 15 min read 8 25 1. Introduction & tl;dr (See changelog at the bottom for some post-publication edits.) 1.1 tl;dr I will argue that a large class of reward functions, which I call “behaviorist”, and which includes almost every reward function in the RL and LLM literature, are all doomed to eventually lead to AI that will “scheme”—i.e., pretend to be docile and cooperative while secretly looking for opport
Explore this link on the map →related reading
- The behavioral selection model for predicting AI motivations — LessWronglesswrong.com
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Reward Function Design: a starter pack — LessWronglesswrong.com
- Reward is not the optimization target — LessWronglesswrong.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Training-time schemers vs behavioral schemers — LessWronglesswrong.com
- Reward hacking is becoming more sophisticated and deliberate in frontier LLMs — LessWronglesswrong.com
- Reward Is Not Enough — LessWronglesswrong.com
- Scheming AIs Will AIs fake alignment during training in order to get power?arxiv.org
- How will we update about scheming? — LessWronglesswrong.com
- Models Don't "Get Reward" — LessWronglesswrong.com