HANABI – np – ( ´ ▽ ` )ノ
nphard.io · 8,589 words · saved by 1 readers
. . . .'. \ / \ / .'. .' '.' ' -= o =- -= o =- .' ' / \ / \ ' '
HANABI . . . .'. \ / \ / .'. .' '.' ' -= o =- -= o =- .' ' / \ / \ ' ' In this post I will go through how I implemented multi-agent environments using Prime Intellect’s stack as part of their RL Residency. My objective is two-fold: To show how multi-agent environments can already be designed using the verifiers library and how training can be done on such environments using both prime-rl and hosted training . To propose and discuss abstractions that could be included into verifiers to allow for more ergonomic multi-agent designs in the future. The main focus is Hanabi , a cooperative card game
related reading
- Natural Deception with RL - Rajan Agarwalrajan.sh
- [2206.12765] Generalized Beliefs for Cooperative AIarxiv.org
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- A Taxonomy of RL Environments for LLM Agentsleehanchung.github.io
- Learning To Play Settlers of Catan With Deep RLsettlers-rl.github.io
- (Some) Natural Emergent Misalignment from Reward Hacking in Non-Production RL — LessWronglesswrong.com
- How training-gamers might function (and win)blog.redwoodresearch.org
- [AN #70]: Agents that help humans who are still learning about their own preferences — LessWronglesswrong.com
- How DeepMind's Generally Capable Agents Were Trained — LessWronglesswrong.com
- combinatorial reasoning environments for LLMs and RLdjdumpling.github.io
- Going Beyond the Nash Equilibrium in Poker by Learning to Exploit Suboptimal Playarxiv.org
- [2608.03958] A game theory for foundation models shows new paths to rational cooperation through similarity inferencearxiv.org