Natural Deception with RL - Rajan Agarwal
rajan.sh · 2,279 words · saved by 2 readers
Language models, when trained on hidden-information games, naturally learn deceptive techniques to win the game by any means.
Natural Deception with RL - Rajan Agarwal TL;DR : We built an RL environment for hidden-information games and assigned binary rewards to win, training Qwen7B to play against other agents. We observed whether the models naturally learned to be deceptive, without intermediate rewards for bluffing or lying. We found that our agent, trained only on win/loss reward, learns strategies that look a lot like deception, such as misreporting their role, faking alliances, and selectively revealing information to manipulate other players. We argue that in multi-turn RL, the frozen agents’ responses act as
saved by
related reading
- Deep Deceptiveness — LessWronglesswrong.com
- How confessions can keep language models honest | OpenAIopenai.com
- Can Agents Fool Each Other? - AI Villagetheaidigest.org
- The Waluigi Effect (mega-post) — LessWronglesswrong.com
- [2602.15515] The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probesarxiv.org
- [2401.05566] Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Trainingarxiv.org
- Models don’t seem to be dishonest in the way humans are — LessWronglesswrong.com
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- On Anthropic's Sleeper Agents Paperthezvi.substack.com
- [2504.04072] Among Us: A Sandbox for Measuring and Detecting Agentic Deceptionarxiv.org
- Emergent Deception and Emergent Optimizationbounded-regret.ghost.io
- HANABI – np – ( ´ ▽ ` )ノnphard.io