flâneur — a map of the web's best reading

Natural Deception with RL - Rajan Agarwal

rajan.sh · 2,279 words · saved by 1 readers

Language models, when trained on hidden-information games, naturally learn deceptive techniques to win the game by any means.

Natural Deception with RL - Rajan Agarwal TL;DR : We built an RL environment for hidden-information games and assigned binary rewards to win, training Qwen7B to play against other agents. We observed whether the models naturally learned to be deceptive, without intermediate rewards for bluffing or lying. We found that our agent, trained only on win/loss reward, learns strategies that look a lot like deception, such as misreporting their role, faking alliances, and selectively revealing information to manipulate other players. We argue that in multi-turn RL, the frozen agents’ responses act as

Explore this link on the map →

saved by

related reading