flâneur — a map of the web's best reading

Rethinking RL Infra for Agents | B'Log

billxbf.github.io · 2,852 words · saved by 1 readers

Why the agentic shift breaks classical RL infra, a tour of Forge, ROLL, SkyRL and Slime, and my recent take with Polar (Agentic RL on Any Harness at Scale).

Table of Contents A year ago, RL for LLMs was almost entirely about reasoning - math, code, short single-turn problems with a clean verifier. The whole loop fit on one picture: model generates a response, environment scores it, optimizer updates the weights. Research focus were algorithmic (GRPO, DAPO, GSPO, Dr.GRPO); while infra was mostly an afterthought. Agentic RL breaks that picture. ROLL puts it well in their report: RLVR trains models that "can answer"; agentic RL trains models that "can act." The model no longer produces a single response - it lives inside a harness, calls tools, reads

Explore this link on the map →

saved by

related reading