[2506.19733] Breaking Barriers: Do Reinforcement Post Training Gains Transfer To Unseen Domains?
Abstract:Reinforcement post training (RPT) has recently shown promise in improving the reasoning abilities of large language models (LLMs). However, it remains unclear how well these improvements generalize to new domains, as prior work evaluates RPT models on data from the same domains used for post-training. To understand the generalizability of RPT, we conduct two studies with specific focus on Reinforcement Learning with Verifiable Rewards (RLVR). (1) Observational: we compare a wide range of open-weight RPT models against their corresponding base models across multiple domains, including both seen and unseen domains in their fine-tuning data. (2) Interventional: we fine-tune LLMs with RPT on single domains and evaluate their performance across multiple domains. Both studies converge on the same conclusion that, although RPT brings substantial gains on tasks similar to the fine-tuning data, the gains generalize inconsistently and can vanish on domains with different reasoning patterns.
# link_2frkvhx09k3.pdf ## Metadata - PDFFormatVersion=1.7 - IsLinearized=false - IsAcroFormPresent=false - IsXFAPresent=false - IsCollectionPresent=false - IsSignaturesPresent=false - Author=Chuxuan Hu; Yuxuan Zhu; Antony Kellermann; Caleb Biddulph; Suppakit Waiwitlikhit; Jason Benn; Daniel Kang - Creator=arXiv GenPDF (tex2pdf:57610bf) - Custom.DOI=https://doi.org/10.48550/arXiv.2506.19733 - Custom.License=http://arxiv.org/licenses/nonexclusive-distrib/1.0/ - Custom.PTEX.Fullbanner=This is pdfTeX, Version 3.141592653-2.6-1.40.28 (TeX Live 2025) kpathsea version 6.4.1 - Custom.arXivID=https://a
Explore this link on the map →saved by
related reading
- DeepSeek-R1arxiv.org
- The State of Reinforcement Learning for LLM Reasoningmagazine.sebastianraschka.com
- State of RL for reasoning LLMs | A. Weersaweers.de
- [2502.19402] General Reasoning Requires Learning to Reason from the Get-goar5iv.labs.arxiv.org
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- Composer2.pdfcursor.com
- Explore | alphaXivalphaxiv.org
- Why reasoning models will generalize - by Nathan Lambertinterconnects.ai
- The State of Reinforcement Learning for LLM Reasoningsebastianraschka.com
- How do LLMs generalize when we do training that is intuitively compatible with two off-distribution behaviors? — LessWronglesswrong.com
- [2602.05910] Chunky Post-Training: Data Driven Failures of Generalizationarxiv.org
- To Understand Language is to Understand Generalization | Eric Jangevjang.com