flâneur — a map of the web's best reading

“Behaviorist” RL reward functions lead to scheming — AI Alignment Forum

alignmentforum.org · 5,900 words · saved by 1 readers

I will argue that a large class of reward functions, which I call “behaviorist”, and which includes almost every reward function in the RL and LLM literature, are all doomed to eventually lead to AI that will “scheme”—i.e., pretend to be docile and cooperative while secretly looking for opportunities to behave in egregiously bad ways such as world takeover (cf. “treacherous turn”). I’ll mostly focus on “brain-like AGI” (as defined just below), but I think the argument applies equally well to future LLMs, if their competence comes overwhelmingly from RL rather than from pretraining.[1] The issue is basically that “negative reward for lying and stealing” looks the same as “negative reward for getting caught lying and stealing”. I’ll argue that the AI will wind up with the latter motivation. The reward function will miss sufficiently sneaky misaligned behavior, and so the AI will come to feel like that kind of behavior is good, and this tendency will generalize in a very bad way. What ver

x “Behaviorist” RL reward functions lead to scheming — AI Alignment Forum Deceptive Alignment AI Frontpage 25 “Behaviorist” RL reward functions lead to scheming by Steven Byrnes 23rd Jul 2025 15 min read 8 25 1. Introduction & tl;dr (See changelog at the bottom for some post-publication edits.) 1.1 tl;dr I will argue that a large class of reward functions, which I call “behaviorist”, and which includes almost every reward function in the RL and LLM literature, are all doomed to eventually lead to AI that will “scheme”—i.e., pretend to be docile and cooperative while secretly looking for opport

Explore this link on the map →

related reading