✳flâneur — a map of the web's best reading
HANABI – np – ( ´ ▽ ` )ノ
nphard.io · 8,589 words · saved by 1 readers
. . . .'. \ / \ / .'. .' '.' ' -= o =- -= o =- .' ' / \ / \ ' '
HANABI . . . .'. \ / \ / .'. .' '.' ' -= o =- -= o =- .' ' / \ / \ ' ' In this post I will go through how I implemented multi-agent environments using Prime Intellect’s stack as part of their RL Residency. My objective is two-fold: To show how multi-agent environments can already be designed using the verifiers library and how training can be done on such environments using both prime-rl and hosted training . To propose and discuss abstractions that could be included into verifiers to allow for more ergonomic multi-agent designs in the future. The main focus is Hanabi , a cooperative card game
Explore this link on the map →related reading
- Natural Deception with RL - Rajan Agarwalrajan.sh
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Composer2.pdfcursor.com
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- [AN #70]: Agents that help humans who are still learning about their own preferences — LessWronglesswrong.com
- How DeepMind's Generally Capable Agents Were Trained — LessWronglesswrong.com
- A Taxonomy of RL Environments for LLM Agentsleehanchung.github.io
- Generally capable agents emerge from open-ended play — Google DeepMinddeepmind.google
- Learning To Play Settlers of Catan With Deep RLsettlers-rl.github.io
- A Toy Environment For Exploring Reasoning About Reward — LessWronglesswrong.com
- VGC-Bench: Towards Mastering Diverse Team Strategies in Competitive Pokémonarxiv.org
- Theory of Mind for Multi-Agent Collaboration via Large Language Modelsarxiv.org