flâneur — a map of the web's best reading

Training Composer for longer horizons · Cursor

cursor.com · 1,675 words · saved by 2 readers

By making self-summarization part of Composer's training, we can get training signal from trajectories much longer than the model's max context window.

Blog / research We train Composer for long-horizon tasks through a reinforcement learning process called self-summarization. By making self-summarization part of Composer's training, we can get training signal from trajectories much longer than the model's max context window. This translates into Composer being able to learn to work on challenging coding tasks requiring hundreds of actions. # The limits of compaction techniques In CursorBench , our internal benchmark suite, we observe that better performance on challenging real-world coding tasks is directly correlated with more thinking and c

Explore this link on the map →

saved by

related reading