✳flâneur — a map of the web's best reading
Natural Deception with RL - Rajan Agarwal
rajan.sh · 2,279 words · saved by 1 readers
Language models, when trained on hidden-information games, naturally learn deceptive techniques to win the game by any means.
Natural Deception with RL - Rajan Agarwal TL;DR : We built an RL environment for hidden-information games and assigned binary rewards to win, training Qwen7B to play against other agents. We observed whether the models naturally learned to be deceptive, without intermediate rewards for bluffing or lying. We found that our agent, trained only on win/loss reward, learns strategies that look a lot like deception, such as misreporting their role, faking alliances, and selectively revealing information to manipulate other players. We argue that in multi-turn RL, the frozen agents’ responses act as
Explore this link on the map →saved by
related reading
- Alignment faking in large language modelsarxiv.org
- How confessions can keep language models honest | OpenAIopenai.com
- Can Agents Fool Each Other? - AI Villagetheaidigest.org
- The Waluigi Effect (mega-post) — LessWronglesswrong.com
- [2401.05566] Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Trainingarxiv.org
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- [2504.04072] Among Us: A Sandbox for Measuring and Detecting Agentic Deceptionarxiv.org
- Emergent Deception and Emergent Optimizationbounded-regret.ghost.io
- HANABI – np – ( ´ ▽ ` )ノnphard.io
- Deep Deceptiveness — LessWronglesswrong.com
- Models don’t seem to be dishonest in the way humans are — LessWronglesswrong.com
- 39 - Evan Hubinger on Model Organisms of Misalignment | AXRP - the AI X-risk Research Podcastaxrp.net