flâneur

Progressive Point Matching

prestonfu.com · 1,653 words · saved by 1 readers

Assigning partial credit to improve reinforcement learning for long-horizon reasoning tasks.

Preston Fu September 2026 Today’s LLMs tackle extremely long-horizon tasks that may run continuously for hours or days. Tasks that take humans days or weeks may require language model trajectories containing millions, or eventually billions, of tokens. These capabilities have been enabled by large-scale reinforcement learning (RL). The standard approach is to sample full trajectories and to assign a sparse outcome reward to the full trajectory – a 0 or 1 based on whether the trajectory was successful. Empirically, this simple approach has demonstrated stable performance improvements at…

saved by

related reading