Rethinking RL Infra for Agents | B'Log
Why the agentic shift breaks classical RL infra, a tour of Forge, ROLL, SkyRL and Slime, and my recent take with Polar (Agentic RL on Any Harness at Scale).
Table of Contents A year ago, RL for LLMs was almost entirely about reasoning - math, code, short single-turn problems with a clean verifier. The whole loop fit on one picture: model generates a response, environment scores it, optimizer updates the weights. Research focus were algorithmic (GRPO, DAPO, GSPO, Dr.GRPO); while infra was mostly an afterthought. Agentic RL breaks that picture. ROLL puts it well in their report: RLVR trains models that "can answer"; agentic RL trains models that "can act." The model no longer produces a single response - it lives inside a harness, calls tools, reads
Explore this link on the map →saved by
related reading
- Forge: Scalable Agent RL Framework and Algorithm - MiniMax News | MiniMaxminimax.io
- Composer2.pdfcursor.com
- Is Frontier Asynchronous RL Solved? — Luke J. Huangluk-huang.github.io
- State of RL for reasoning LLMs | A. Weersaweers.de
- A Taxonomy of RL Environments for LLM Agentsleehanchung.github.io
- Building Effective AI Agents \ Anthropicanthropic.com
- [2603.21972] Demystifying Reinforcement Learning for Long-Horizon Tool-Using Agents: A Comprehensive Recipearxiv.org
- Infini-AI-Lab on X: "We’re excited to release 𝐀𝐬𝐭𝐫𝐚𝐅𝐥𝐨𝐰, an open-source, dataflow-oriented RL system for training multi-agentic and multi-policy LLMs. 🚀 Built for scalable, flexible, and efficient agent RL, AstraFlow natively enables: ⚡ 𝟐.𝟕× 𝐟𝐚𝐬𝐭𝐞𝐫 𝐦𝐮𝐥𝐭𝐢-𝐩𝐨𝐥𝐢𝐜𝐲 https://t.co/JVthM8iHur" / Xx.com
- Building Effective AI Agents \ Anthropicanthropic.com
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- A vision researcher’s guide to some RL stuff: PPO & GRPO - Yuge (Jimmy) Shiyugeten.github.io
- RL at 1T Scale: prime-rl Performance Deep Diveprimeintellect.ai