Training Math Reasoning Model with Reinforcement Learning - NVIDIA ADLR
We are excited to introduce AceMath-RL-Nemotron-7B, a math model trained purely with reinforcement learning (RL) from the Deepseek-R1-Distilled-Qwen-7B checkpoint. It achieves 69.0% Pass@1 accuracy on AIME 2024 (+12.5% gain from DeepSeek-R1-Distill), surpassing o3-mini (low) at 60% and o1-mini at 63.6%. On AIME 2025, it achieves 53.6% Pass@1 accuracy (+13% gain). Surprisingly, math-focused RL training also improves the model’s coding accuracy on LiveCodeBench, reaching 44.4% Pass@1 (+6.8% gain), demonstrating the generalization capabilities of scaled RL training. We find that using on-policy GRPO training is critical to preventing entropy collapse during training while ensuring sufficient exploration. Curriculum learning: Response length extension training from 8K —> 16K —> 24K —> 32K We evaluate our model against competitive reasoning models of comparable size on AIME 2024, AIME 2025, and GPQA. Additionally, we evaluate our models on additional math benchmarks and LiveCodeBench for a
AceMath-RL-Nemotron AceMath-RL-Nemotron-8B: [Checkpoints🤗] Author: Yang Chen, Zihan Liu, Chankyu Lee, Mohammad Shoeybi, Bryan Catanzaro, Wei Ping Overview We are excited to introduce AceMath-RL-Nemotron-7B , a math model trained purely with reinforcement learning (RL) from the Deepseek-R1-Distilled-Qwen-7B checkpoint. It achieves 69.0% Pass@1 accuracy on AIME 2024 (+12.5% gain from DeepSeek-R1-Distill), surpassing o3-mini (low) at 60% and o1-mini at 63.6%. On AIME 2025, it achieves 53.6% Pass@1 accuracy (+13% gain). Surprisingly, math-focused RL training also improves the model’s coding accur
Explore this link on the map →saved by
related reading
- How We Build Trillion Parameter Reasoning RL with 10% GPUsmacaron.im
- DeepSeek-R1arxiv.org
- The State of Reinforcement Learning for LLM Reasoningmagazine.sebastianraschka.com
- Composer2.pdfcursor.com
- State of RL for reasoning LLMs | A. Weersaweers.de
- Teaching a Language Model Arithmetic with Reinforcement Learning - Sami Khansamikhan.ai
- Explore | alphaXivalphaxiv.org
- Understanding R1-Zero-Like Training: A Critical Perspectivearxiv.org
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- The State of Reinforcement Learning for LLM Reasoningsebastianraschka.com
- the-illusion-of-thinking.pdfml-site.cdn-apple.com
- DeepSeek R1's recipe to replicate o1 and the future of reasoning LMsinterconnects.ai