flâneur — a map of the web's best reading

HANABI – np – ( ´ ▽ ` )ノ

nphard.io · 8,589 words · saved by 1 readers

. . . .'. \ / \ / .'. .' '.' ' -= o =- -= o =- .' ' / \ / \ ' '

HANABI . . . .'. \ / \ / .'. .' '.' ' -= o =- -= o =- .' ' / \ / \ ' ' In this post I will go through how I implemented multi-agent environments using Prime Intellect’s stack as part of their RL Residency. My objective is two-fold: To show how multi-agent environments can already be designed using the verifiers library and how training can be done on such environments using both prime-rl and hosted training . To propose and discuss abstractions that could be included into verifiers to allow for more ergonomic multi-agent designs in the future. The main focus is Hanabi , a cooperative card game

Explore this link on the map →

related reading