flâneur — a map of the web's best reading

Is Frontier Asynchronous RL Solved? — Luke J. Huang

luk-huang.github.io · 3,679 words · saved by 4 readers

A blog post by Luke J. Huang on whether frontier asynchronous RL is solved, covering policy lag, stability methods, open questions, and appendix notes.

Table of Contents Async RL has become the default for large-scale RL post-training. Frontier open-weights labs — GLM-5 , Ring 1T , DeepSeek V3.2 , Minimax M2.5 , Qwen 3.5 , Intellect-3 , Nemotron-3 Super , and Laguna-M.1 — report 2–3× faster throughput over synchronous pipelines, each with its own approach to keeping training stable. This is an attempt to survey that landscape: what does each lab do? What are the shared failure modes? Where do things currently stand? TL;DR async RL decouples rollout and training, giving 2-3x throughput, but the stale trajectories create off-policy instability

Explore this link on the map →

saved by

related reading