Improving Composer through real-time RL · Cursor
We apply online reinforcement learning to Composer, serving model checkpoints to production and using real user interactions as reward signals to ship an improved checkpoint multiple times a day.
Blog / research We are observing unprecedented growth in the usefulness and adoption of coding models in the real world. In the face of 10–100x increases in inference volume, we consider the question: how can we take these trillions of tokens and extract from them a training signal to improve the model? We call our approach of using real inference tokens for training "real-time RL." We first used this technique to train Tab and we found it was highly effective. Now we're applying a similar approach to Composer. We serve model checkpoints to production, observe user responses, and aggregate tho
saved by
related reading
- Composer2.pdfcursor.com
- Training Composer for longer horizons · Cursorcursor.com
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- Improving Cursor Tab with online RL · Cursorcursor.com
- Coding vs thinking — Paradigm 3paradigm3.org
- Simulated Users & Sad LLMs1a3orn.com
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Sonya Huang 🐥 on X: "Today's Training Data episode takes us BTS on the infrastructure challenges required to do large RL runs at scale, featuring @ellev3n11 (Composer Lead at @cursor_ai) and @dzhulgakov (Co-Founder at @FireworksAI_HQ). The Cursor team trained Composer 2 on Fireworks by starting with https://t.co/6LLlJlyl8Q" / Xx.com
- Kevin-32B: Multi-Turn RL for Writing CUDA Kernels | Cognitioncognition.ai
- [2607.18966] Measuring Reward-Seeking via Contrastive Belief Updatesarxiv.org
- Unlocking Real-Time Bug Detection at Cognition | Applied Computeappliedcompute.com