Training Composer for longer horizons · Cursor
By making self-summarization part of Composer's training, we can get training signal from trajectories much longer than the model's max context window.
Blog / research We train Composer for long-horizon tasks through a reinforcement learning process called self-summarization. By making self-summarization part of Composer's training, we can get training signal from trajectories much longer than the model's max context window. This translates into Composer being able to learn to work on challenging coding tasks requiring hundreds of actions. # The limits of compaction techniques In CursorBench , our internal benchmark suite, we observe that better performance on challenging real-world coding tasks is directly correlated with more thinking and c
Explore this link on the map →saved by
related reading
- Composer2.pdfcursor.com
- Improving Composer through real-time RL · Cursorcursor.com
- Effective harnesses for long-running agents \ Anthropicanthropic.com
- The First Fully General Computer Action Model | blogsi.inc
- LLM Powered Autonomous Agents | Lil'Loglilianweng.github.io
- Effective context engineering for AI agents \ Anthropicanthropic.com
- LLMs can invent their own compression - Rajan Agarwalrajan.sh
- Why We Need Continual Learning | Andreessen Horowitza16z.com
- Shipping at Inference-Speed | Peter Steinbergersteipete.me
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- [2506.06266] Cartridges: Lightweight and general-purpose long context representations via self-studyarxiv.org
- Coding vs thinking — Paradigm 3paradigm3.org