“Behaviorist” RL reward functions lead to scheming — AI Alignment Forum
I will argue that a large class of reward functions, which I call “behaviorist”, and which includes almost every reward function in the RL and LLM literature, are all doomed to eventually lead to AI that will “scheme”—i.e., pretend to be docile and cooperative while secretly looking for opportunities to behave in egregiously bad ways such as world takeover (cf. “treacherous turn”). I’ll mostly focus on “brain-like AGI” (as defined just below), but I think the argument applies equally well to future LLMs, if their competence comes overwhelmingly from RL rather than from pretraining.[1] The issue is basically that “negative reward for lying and stealing” looks the same as “negative reward for getting caught lying and stealing”. I’ll argue that the AI will wind up with the latter motivation. The reward function will miss sufficiently sneaky misaligned behavior, and so the AI will come to feel like that kind of behavior is good, and this tendency will generalize in a very bad way. What ver
x “Behaviorist” RL reward functions lead to scheming — AI Alignment Forum Deceptive Alignment AI Frontpage 25 “Behaviorist” RL reward functions lead to scheming by Steven Byrnes 23rd Jul 2025 15 min read 8 25 1. Introduction & tl;dr (See changelog at the bottom for some post-publication edits.) 1.1 tl;dr I will argue that a large class of reward functions, which I call “behaviorist”, and which includes almost every reward function in the RL and LLM literature, are all doomed to eventually lead to AI that will “scheme”—i.e., pretend to be docile and cooperative while secretly looking for opport
related reading
- The behavioral selection model for predicting AI motivations — LessWronglesswrong.com
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- New report: “Scheming AIs: Will AIs fake alignment during training in order to get power?”joecarlsmith.com
- Reward Function Design: a starter pack — LessWronglesswrong.com
- Reward is not the optimization target — LessWronglesswrong.com
- Many arguments for AI x-risk are wrong — AI Alignment Forumalignmentforum.org
- Training a Misaligned Reward Seekeralignment.anthropic.com
- Why are AI agents lying, cheating and coordinating?yoshuabengio.org
- Reward hacking is becoming more sophisticated and deliberate in frontier LLMs — LessWronglesswrong.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Training-time schemers vs behavioral schemers — LessWronglesswrong.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com