flâneur — a map of the web's best reading

Why do Policy Gradient Methods work so well in Cooperative MARL? Evidence from Policy Representation – The Berkeley Artificial Intelligence Research Blog

bair.berkeley.edu · 1,357 words · saved by 1 readers

The BAIR Blog

In cooperative multi-agent reinforcement learning (MARL), due to its on-policy nature, policy gradient (PG) methods are typically believed to be less sample efficient than value decomposition (VD) methods, which are off-policy . However, some recent empirical studies demonstrate that with proper input representation and hyper-parameter tuning, multi-agent PG can achieve surprisingly strong performance compared to off-policy VD methods. Why could PG methods work so well? In this post, we will present concrete analysis to show that in certain scenarios, e.g., environments with a highly multi-mod

Explore this link on the map →

related reading