flâneur — a map of the web's best reading

You Only Need Minimal RLVR Training: Extrapolating LLMs via Rank-1 Trajectories

arxiv.org · 8,624 words · saved by 1 readers

Reinforcement learning with verifiable rewards (RLVR) has become a dominant paradigm for improving reasoning in large language models (LLMs), yet the underlying geometry of the resulting parameter trajectories remains underexplored. In this work, we demonstrate that RLVR weight trajectories are extremely low-rank and highly predictable. Specifically, we find that the majority of downstream performance gains are captured by a rank-1 approximation of the parameter deltas, where the magnitude of this projection evolves near-linearly with training steps. Motivated by this, we propose a simple and compute-efficient method RELEX (REinforcement Learning EXtrapolation), which estimates the rank-1 subspace from a short observation window and extrapolates future checkpoints via linear regression, with no learned model required. Across three models (i.e., Qwen2.5-Math-1.5B, Qwen3-4B-Base, and Qwen3-8B-Base), RELEX produces checkpoints that match or exceed RLVR performance on both in-domain and ou

You Only Need Minimal RLVR Training: Extrapolating LLMs via Rank-1 Trajectories Zhepei Wei † Xinyu Zhu † Wei-Lin Chen † Chengsong Huang ‡ Jiaxin Huang ‡ Yu Meng † † University of Virginia ‡ Washington University in St. Louis {zhepei.wei,xinyuzhu,wlchen,yumeng5}@virginia.edu {chengsong,jiaxinh}@wustl.edu Abstract Reinforcement learning with verifiable rewards (RLVR) has become a dominant paradigm for improving reasoning in large language models (LLMs), yet the underlying geometry of the resulting parameter trajectories remains underexplored. In this work, we demonstrate that RLVR weight traject

Explore this link on the map →

saved by

related reading