JaxMARL: Multi-Agent RL, but 10000x Faster
We present JaxMARL, a library of multi-agent reinforcement learning (MARL) environments and algorithms based on end-to-end GPU acceleration that achieves up to 12500x speedups. The environments in JaxMARL span cooperative, competitive, and mixed games; discrete and continuous state and action spaces; and zero-shot and CTDE settings. We specifically include implementations of the Hanabi Learning Environment, Overcooked, Multi-Agent Brax, MPE, Switch Riddle, Coin Game, and Spatial-Temporal Representations of Matrix Games (STORM). Because of JAX's hardware acceleration, our per-run training pipeline is 12500x faster than existing approaches. We also introduce SMAX, a vectorised version of the popular StarCraft Multi-Agent Challenge, which removes the need to run the StarCraft II game engine. By significantly speeding up training, JaxMARL enables new research into areas such as multi-agent meta-learning, as well as significantly easing and improving evaluation in MARL. Try it out here: htt
JaxMARL: Multi-Agent RL, but 10000x Faster Overview We present JaxMARL , a library of multi-agent reinforcement learning (MARL) environments and algorithms based on end-to-end GPU acceleration that achieves up to 12500x speedups. The environments in JaxMARL span cooperative, competitive, and mixed games; discrete and continuous state and action spaces; and zero-shot and CTDE settings. We specifically include implementations of the Hanabi Learning Environment, Overcooked, Multi-Agent Brax, MPE, Switch Riddle, Coin Game, and Spatial-Temporal Representations of Matrix Games (STORM). Because of JA
Explore this link on the map →related reading
- GitHub - FLAIROx/JaxMARL: Multi-Agent Reinforcement Learning with JAX · GitHubgithub.com
- MARLlib: A Multi-agent Reinforcement Learning Library — MARLlib v1.0.0 documentationmarllib.readthedocs.io
- GitHub - marlbenchmark/on-policy: This is the official implementation of Multi-Agent PPO (MAPPO). · GitHubgithub.com
- The 37 Implementation Details of Proximal Policy Optimization · The ICLR Blog Trackiclr-blog-track.github.io
- A Taxonomy of RL Environments for LLM Agentsleehanchung.github.io
- Forge: Scalable Agent RL Framework and Algorithm - MiniMax News | MiniMaxminimax.io
- Multi-Agent Reinforcement Learning: Foundations and Modern Approachesmarl-book.com
- craftaxenv.github.iocraftaxenv.github.io
- Why do Policy Gradient Methods work so well in Cooperative MARL? Evidence from Policy Representation – The Berkeley Artificial Intelligence Research Blogbair.berkeley.edu
- [2605.22748] Superhuman Safe and Agile Racing through Multi-Agent Reinforcement Learningarxiv.org
- How DeepMind's Generally Capable Agents Were Trained — LessWronglesswrong.com
- Infini-AI-Lab on X: "We’re excited to release 𝐀𝐬𝐭𝐫𝐚𝐅𝐥𝐨𝐰, an open-source, dataflow-oriented RL system for training multi-agentic and multi-policy LLMs. 🚀 Built for scalable, flexible, and efficient agent RL, AstraFlow natively enables: ⚡ 𝟐.𝟕× 𝐟𝐚𝐬𝐭𝐞𝐫 𝐦𝐮𝐥𝐭𝐢-𝐩𝐨𝐥𝐢𝐜𝐲 https://t.co/JVthM8iHur" / Xx.com