Sonya Huang 🐥 on X: "Today's Training Data episode takes us BTS on the infrastructure challenges required to do large RL runs at scale, featuring @ellev3n11 (Composer Lead at @cursor_ai) and @dzhulgakov (Co-Founder at @FireworksAI_HQ). The Cursor team trained Composer 2 on Fireworks by starting with https://t.co/6LLlJlyl8Q" / X
To view keyboard shortcuts, press question mark View keyboard shortcuts Home Explore Notifications Chat Grok Premium Money Bookmarks Creator Studio Articles Profile More Post christina @luoluo Post See new posts Conversation Federico Cassano reposted Sonya Huang @sonyatweetybird Today's Training Data episode takes us BTS on the infrastructure challenges required to do large RL runs at scale, featuring @ellev3n11 (Composer Lead at @cursor_ai ) and @dzhulgakov (Co-Founder at @FireworksAI_HQ ). The Cursor team trained Composer 2 on Fireworks by starting with a strong base model (Kimi 2.5) and performing large-scale mid-training on code tokens and web data to learn common patterns and libraries, followed by a large-scale Reinforcement Learning run to learn how to navigate the Cursor harness, call tools, and write correct code. Today's episode dives into the systems and infrastructure challenges of making that large RL run happening, and there were many (!!), from numerical mismat
@sonyatweetybird: Today's Training Data episode takes us BTS on the infrastructure challenges required to do large RL runs at scale, featuring @ellev3n11 (Composer Lead at @cursor_ai ) and @dzhulgakov (Co-Founder at @FireworksAI_HQ ). The Cursor team trained Composer 2 on Fireworks by starting with a strong base model (Kimi 2.5) and performing large-scale mid-training on code tokens and web data to learn common patterns and libraries, followed by a large-scale Reinforcement Learning run to learn how to navigate the Cursor harness, call tools, and write correct code. Today's episode di
Explore this link on the map →saved by
related reading
- Infini-AI-Lab on X: "We’re excited to release 𝐀𝐬𝐭𝐫𝐚𝐅𝐥𝐨𝐰, an open-source, dataflow-oriented RL system for training multi-agentic and multi-policy LLMs. 🚀 Built for scalable, flexible, and efficient agent RL, AstraFlow natively enables: ⚡ 𝟐.𝟕× 𝐟𝐚𝐬𝐭𝐞𝐫 𝐦𝐮𝐥𝐭𝐢-𝐩𝐨𝐥𝐢𝐜𝐲 https://t.co/JVthM8iHur" / Xx.com
- Composer2.pdfcursor.com
- Is Frontier Asynchronous RL Solved? — Luke J. Huangluk-huang.github.io
- Improving Composer through real-time RL · Cursorcursor.com
- Training great LLMs entirely from ground up in the wilderness as a startup - Yi Tayyitay.net
- IsoCompute Playbook: Optimally Scaling Sampling Compute for RL Training of LLMscompute-optimal-rl-llm-scaling.github.io
- Reinforcement learning is an infrastructure problemmodal.com
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- RL Scaling Laws for LLMs - by Cameron R. Wolfe, Ph.D.cameronrwolfe.substack.com
- Orbit - Ultra-efficient RL Pipelinespherelab.ai
- [2603.21972] Demystifying Reinforcement Learning for Long-Horizon Tool-Using Agents: A Comprehensive Recipearxiv.org
- State of RL for reasoning LLMs | A. Weersaweers.de