Wang: Loongrl: Reinforcement learning for advanced... - Google Scholar
Articles 6 results (0.02 sec) My profile My library Any time Since 2026 Since 2025 Since 2022 Custom range... Citations 2025 2026 Sort by relevance Sort by date Create alert Loongrl: Reinforcement learning for advanced reasoning over long contexts Search within citing articles [PDF] arxiv.org Evidence-Augmented Policy Optimization with Reward Co-Evolution for Long-Context Reasoning X Guan, Z Li, S Huang, P Xie, J Zhou, J Cao - arXiv preprint arXiv …, 2026 - arxiv.org While Reinforcement Learning (RL) has advanced LLM reasoning, applying it to long- context scenarios is hindered by sparsity of outcome rewards. This limitation fails to penalize … Save Cite Cited by 1 Related articles All 2 versions [PDF] arxiv.org LongAct: Harnessing Intrinsic Activation Patterns for Long-Context Reinforcement Learning B Ping, Z Chen, T Hui, Q Yu, C Li, J Yan… - arXiv preprint arXiv …, 2026 - arxiv.org Reinforcement Learning (RL) has emerged as a critical driver for enhancing the reasoning capabilitie
[PDF] arxiv.org Beyond Teacher Likelihood: Group-Calibrated On-Policy Distillation for Long-Context Reasoning Z Zhang, J Wang, X Xu, X Wang, Z Zhou…�- arXiv preprint arXiv�…, 2026 - arxiv.org On-policy distillation (OPD) trains a student on its own responses using dense token-level guidance from a stronger teacher. In long-context tasks, however, token-level teacher�… Save Cite Cited by 1 Related articles View as HTML [PDF] arxiv.org Tracking the Moving Frontier: Long-Short Term Advantage Estimator X Yao, L Yu, C Wang, F Teng, Y Zhang, Q Cui…�- arXiv preprint arXiv�…, 2026 -…
saved by
related reading
- QwenLong-L1.5: Post-Training Recipe for Long-Context Reasoning and Memory Managementarxiv.org
- State of RL for reasoning LLMs | A. Weersaweers.de
- DeepSeek-R1arxiv.org
- On-Policy Distillation - Thinking Machines Labthinkingmachines.ai
- The State of Reinforcement Learning for LLM Reasoningmagazine.sebastianraschka.com
- Recursive Language Models | Alex L. Zhangalexzhang13.github.io
- Explore | alphaXivalphaxiv.org
- GRPO++: Tricks for Making RL Actually Workcameronrwolfe.substack.com
- Extending Context is Hard | kaiokendevkaiokendev.github.io
- The State of Reinforcement Learning for LLM Reasoningsebastianraschka.com
- The huge potential implications of long-context inferenceepochai.substack.com
- Progressive Point Matchingprestonfu.com