flâneur — a map of the web's best reading

Training Math Reasoning Model with Reinforcement Learning - NVIDIA ADLR

research.nvidia.com · 1,075 words · saved by 1 readers

We are excited to introduce AceMath-RL-Nemotron-7B, a math model trained purely with reinforcement learning (RL) from the Deepseek-R1-Distilled-Qwen-7B checkpoint. It achieves 69.0% Pass@1 accuracy on AIME 2024 (+12.5% gain from DeepSeek-R1-Distill), surpassing o3-mini (low) at 60% and o1-mini at 63.6%. On AIME 2025, it achieves 53.6% Pass@1 accuracy (+13% gain). Surprisingly, math-focused RL training also improves the model’s coding accuracy on LiveCodeBench, reaching 44.4% Pass@1 (+6.8% gain), demonstrating the generalization capabilities of scaled RL training. We find that using on-policy GRPO training is critical to preventing entropy collapse during training while ensuring sufficient exploration. Curriculum learning: Response length extension training from 8K —> 16K —> 24K —> 32K We evaluate our model against competitive reasoning models of comparable size on AIME 2024, AIME 2025, and GPQA. Additionally, we evaluate our models on additional math benchmarks and LiveCodeBench for a

AceMath-RL-Nemotron AceMath-RL-Nemotron-8B: [Checkpoints🤗] Author: Yang Chen, Zihan Liu, Chankyu Lee, Mohammad Shoeybi, Bryan Catanzaro, Wei Ping Overview We are excited to introduce AceMath-RL-Nemotron-7B , a math model trained purely with reinforcement learning (RL) from the Deepseek-R1-Distilled-Qwen-7B checkpoint. It achieves 69.0% Pass@1 accuracy on AIME 2024 (+12.5% gain from DeepSeek-R1-Distill), surpassing o3-mini (low) at 60% and o1-mini at 63.6%. On AIME 2025, it achieves 53.6% Pass@1 accuracy (+13% gain). Surprisingly, math-focused RL training also improves the model’s coding accur

Explore this link on the map →

saved by

related reading