Sonya Huang 🐥 on X: "Today's Training Data episode takes us BTS on the infrastructure challenges required to do large RL runs at scale, featuring @ellev3n11 (Composer Lead at @cursor_ai) and @dzhulgakov (Co-Founder at @FireworksAI_HQ). The Cursor team trained Composer 2 on Fireworks by starting with https://t.co/6LLlJlyl8Q" / X
To view keyboard shortcuts, press question mark View keyboard shortcuts Home Explore Notifications Chat Grok Premium Money Bookmarks Creator Studio Articles Profile More Post christina @luoluo Post See new posts Conversation Federico Cassano reposted Sonya Huang @sonyatweetybird Today's Training Data episode takes us BTS on the infrastructure challenges required to do large RL runs at scale, featuring @ellev3n11 (Composer Lead at @cursor_ai ) and @dzhulgakov (Co-Founder at @FireworksAI_HQ ). The Cursor team trained Composer 2 on Fireworks by starting with a strong base model (Kimi 2.5) and performing large-scale mid-training on code tokens and web data to learn common patterns and libraries, followed by a large-scale Reinforcement Learning run to learn how to navigate the Cursor harness, call tools, and write correct code. Today's episode dives into the systems and infrastructure challenges of making that large RL run happening, and there were many (!!), from numerical mismat
@sonyatweetybird: Today's Training Data episode takes us BTS on the infrastructure challenges required to do large RL runs at scale, featuring @ellev3n11 (Composer Lead at @cursor_ai ) and @dzhulgakov (Co-Founder at @FireworksAI_HQ ). The Cursor team trained Composer 2 on Fireworks by starting with a strong base model (Kimi 2.5) and performing large-scale mid-training on code tokens and web data to learn common patterns and libraries, followed by a large-scale Reinforcement Learning run to learn how to navigate the Cursor harness, call tools, and write correct code. Today's episode di
saved by
related reading
- RL at 1T Scale: prime-rl Performance Deep Diveprimeintellect.ai
- Composer2.pdfcursor.com
- Is Frontier Asynchronous RL Solved? — Luke J. Huangluk-huang.github.io
- Improving Composer through real-time RL · Cursorcursor.com
- The upcoming GPT-3 moment for RL | Mechanize, Inc.mechanize.work
- Training great LLMs entirely from ground up in the wilderness as a startup - Yi Tayyitay.net
- The Extreme Inefficiency of RL for Frontier Models - Toby Ordtobyord.com
- Lee Robinson (@leerob) on Xx.com
- IsoCompute Playbook: Optimally Scaling Sampling Compute for RL Training of LLMscompute-optimal-rl-llm-scaling.github.io
- RL Post-Training on Macs | Pluralis Researchpluralis.ai
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- Reinforcement learning is an infrastructure problemmodal.com