Improving Composer through real-time RL · Cursor
We apply online reinforcement learning to Composer, serving model checkpoints to production and using real user interactions as reward signals to ship an improved checkpoint multiple times a day.
Blog / research We are observing unprecedented growth in the usefulness and adoption of coding models in the real world. In the face of 10–100x increases in inference volume, we consider the question: how can we take these trillions of tokens and extract from them a training signal to improve the model? We call our approach of using real inference tokens for training "real-time RL." We first used this technique to train Tab and we found it was highly effective. Now we're applying a similar approach to Composer. We serve model checkpoints to production, observe user responses, and aggregate tho
Explore this link on the map →saved by
related reading
- Composer2.pdfcursor.com
- Training Composer for longer horizons · Cursorcursor.com
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- Interaction Models: A Scalable Approach to Human-AI Collaboration - Thinking Machines Labthinkingmachines.ai
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- Improving Cursor Tab with online RL · Cursorcursor.com
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Sonya Huang 🐥 on X: "Today's Training Data episode takes us BTS on the infrastructure challenges required to do large RL runs at scale, featuring @ellev3n11 (Composer Lead at @cursor_ai) and @dzhulgakov (Co-Founder at @FireworksAI_HQ). The Cursor team trained Composer 2 on Fireworks by starting with https://t.co/6LLlJlyl8Q" / Xx.com
- Kevin-32B: Multi-Turn RL for Writing CUDA Kernels | Cognitioncognition.ai
- Adithya S K on X: "RL Coding Environments 101: Why Harbor Exists" / Xx.com
- Scaling Reinforcement Learning: Environments, Reward Hacking, Agents, Scaling Datasemianalysis.com
- Coding vs thinking — Paradigm 3paradigm3.org