flâneur — a map of the web's best reading

Wang: Loongrl: Reinforcement learning for advanced... - Google Scholar

scholar.google.com · saved by 1 readers

Articles 6 results (0.02 sec) My profile My library Any time Since 2026 Since 2025 Since 2022 Custom range... Citations 2025 2026 Sort by relevance Sort by date Create alert Loongrl: Reinforcement learning for advanced reasoning over long contexts Search within citing articles [PDF] arxiv.org Evidence-Augmented Policy Optimization with Reward Co-Evolution for Long-Context Reasoning X Guan, Z Li, S Huang, P Xie, J Zhou, J Cao - arXiv preprint arXiv …, 2026 - arxiv.org While Reinforcement Learning (RL) has advanced LLM reasoning, applying it to long- context scenarios is hindered by sparsity of outcome rewards. This limitation fails to penalize … Save Cite Cited by 1 Related articles All 2 versions [PDF] arxiv.org LongAct: Harnessing Intrinsic Activation Patterns for Long-Context Reinforcement Learning B Ping, Z Chen, T Hui, Q Yu, C Li, J Yan… - arXiv preprint arXiv …, 2026 - arxiv.org Reinforcement Learning (RL) has emerged as a critical driver for enhancing the reasoning capabilitie

Articles 6 results (0.02 sec) My profile My library Any time Since 2026 Since 2025 Since 2022 Custom range... Citations 2025 2026 Sort by relevance Sort by date Create alert Loongrl: Reinforcement learning for advanced reasoning over long contexts Search within citing articles [PDF] arxiv.org Evidence-Augmented Policy Optimization with Reward Co-Evolution for Long-Context Reasoning X Guan, Z Li, S Huang, P Xie, J Zhou, J Cao - arXiv preprint arXiv …, 2026 - arxiv.org While Reinforcement Learning (RL) has advanced LLM reasoning, applying it to long- context scenarios is hindered by sparsity o

Explore this link on the map →

saved by