Is Frontier Asynchronous RL Solved? — Luke J. Huang
A blog post by Luke J. Huang on whether frontier asynchronous RL is solved, covering policy lag, stability methods, open questions, and appendix notes.
Table of Contents Async RL has become the default for large-scale RL post-training. Frontier open-weights labs — GLM-5 , Ring 1T , DeepSeek V3.2 , Minimax M2.5 , Qwen 3.5 , Intellect-3 , Nemotron-3 Super , and Laguna-M.1 — report 2–3× faster throughput over synchronous pipelines, each with its own approach to keeping training stable. This is an attempt to survey that landscape: what does each lab do? What are the shared failure modes? Where do things currently stand? TL;DR async RL decouples rollout and training, giving 2-3x throughput, but the stale trajectories create off-policy instability
Explore this link on the map →saved by
related reading
- RL Scaling Laws for LLMs - by Cameron R. Wolfe, Ph.D.cameronrwolfe.substack.com
- State of RL for reasoning LLMs | A. Weersaweers.de
- Infini-AI-Lab on X: "We’re excited to release 𝐀𝐬𝐭𝐫𝐚𝐅𝐥𝐨𝐰, an open-source, dataflow-oriented RL system for training multi-agentic and multi-policy LLMs. 🚀 Built for scalable, flexible, and efficient agent RL, AstraFlow natively enables: ⚡ 𝟐.𝟕× 𝐟𝐚𝐬𝐭𝐞𝐫 𝐦𝐮𝐥𝐭𝐢-𝐩𝐨𝐥𝐢𝐜𝐲 https://t.co/JVthM8iHur" / Xx.com
- Sonya Huang 🐥 on X: "Today's Training Data episode takes us BTS on the infrastructure challenges required to do large RL runs at scale, featuring @ellev3n11 (Composer Lead at @cursor_ai) and @dzhulgakov (Co-Founder at @FireworksAI_HQ). The Cursor team trained Composer 2 on Fireworks by starting with https://t.co/6LLlJlyl8Q" / Xx.com
- Xiuyu Li on X: "RL Interview Questions 2026" / Xx.com
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- Rethinking RL Infra for Agents | B'Logbillxbf.github.io
- Learning Beyond Gradientstrinkle23897.github.io
- Defeating Nondeterminism in LLM Inference - Thinking Machines Labthinkingmachines.ai
- A vision researcher’s guide to some RL stuff: PPO & GRPO - Yuge (Jimmy) Shiyugeten.github.io
- IsoCompute Playbook: Optimally Scaling Sampling Compute for RL Training of LLMscompute-optimal-rl-llm-scaling.github.io
- Composer2.pdfcursor.com