Training Composer for longer horizons · Cursor
By making self-summarization part of Composer's training, we can get training signal from trajectories much longer than the model's max context window.
Blog / research We train Composer for long-horizon tasks through a reinforcement learning process called self-summarization. By making self-summarization part of Composer's training, we can get training signal from trajectories much longer than the model's max context window. This translates into Composer being able to learn to work on challenging coding tasks requiring hundreds of actions. # The limits of compaction techniques In CursorBench , our internal benchmark suite, we observe that better performance on challenging real-world coding tasks is directly correlated with more thinking and c
saved by
related reading
- Coding vs thinking — Paradigm 3paradigm3.org
- Composer2.pdfcursor.com
- Improving Composer through real-time RL · Cursorcursor.com
- SWE-1.7: Frontier Intelligence at a Fraction of the Costcognition.com
- LLMs can invent their own compression - Rajan Agarwalrajan.sh
- The First Fully General Computer Action Model | blogsi.inc
- Effective harnesses for long-running agents \ Anthropicanthropic.com
- Trending Papers - Hugging Facepaperswithcode.com
- GLM-5.2: Built for Long-Horizon Tasksz.ai
- Harness Engineering for Self-Improvement | Lil'Loglilianweng.github.io
- Why We Need Continual Learning | Andreessen Horowitza16z.com
- Shipping at Inference-Speed | Peter Steinbergersteipete.me