Rethinking RL Infra for Agents | B'Log
Why the agentic shift breaks classical RL infra, a tour of Forge, ROLL, SkyRL and Slime, and my recent take with Polar (Agentic RL on Any Harness at Scale).
Table of Contents A year ago, RL for LLMs was almost entirely about reasoning - math, code, short single-turn problems with a clean verifier. The whole loop fit on one picture: model generates a response, environment scores it, optimizer updates the weights. Research focus were algorithmic (GRPO, DAPO, GSPO, Dr.GRPO); while infra was mostly an afterthought. Agentic RL breaks that picture. ROLL puts it well in their report: RLVR trains models that "can answer"; agentic RL trains models that "can act." The model no longer produces a single response - it lives inside a harness, calls tools, reads
saved by
related reading
- Composer2.pdfcursor.com
- How can LLM RL Work Despite Information-Theoretic Inefficiencyberen.io
- Keep the Tokens Flowing: Lessons from 16 Open-Source RL Librarieshuggingface.co
- Is Frontier Asynchronous RL Solved? — Luke J. Huangluk-huang.github.io
- Forge: Scalable Agent RL Framework and Algorithm - MiniMax News | MiniMaxminimax.io
- A Taxonomy of RL Environments for LLM Agentsleehanchung.github.io
- State of RL for reasoning LLMs | A. Weersaweers.de
- The Extreme Inefficiency of RL for Frontier Models - Toby Ordtobyord.com
- RL at 1T Scale: prime-rl Performance Deep Diveprimeintellect.ai
- Building Effective AI Agents \ Anthropicanthropic.com
- Building Effective AI Agents \ Anthropicanthropic.com
- The upcoming GPT-3 moment for RL | Mechanize, Inc.mechanize.work