[2402.03300] DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Abstract:Mathematical reasoning poses a significant challenge for language models due to its complex and structured nature. In this paper, we introduce DeepSeekMath 7B, which continues pre-training DeepSeek-Coder-Base-v1.5 7B with 120B math-related tokens sourced from Common Crawl, together with natural language and code data. DeepSeekMath 7B has achieved an impressive score of 51.7% on the competition-level MATH benchmark without relying on external toolkits and voting techniques, approaching the performance level of Gemini-Ultra and GPT-4. Self-consistency over 64 samples from DeepSeekMath 7B achieves 60.9% on MATH. The mathematical reasoning capability of DeepSeekMath is attributed to two key factors: First, we harness the significant potential of publicly available web data through a meticulously engineered data selection pipeline. Second, we introduce Group Relative Policy Optimization (GRPO), a variant of Proximal Policy Optimization (PPO), that enhances mathematical reasoning abilities while concurrently optimizing the memory usage of PPO.
Authors:Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, Y.K. Li, Y. Wu, Daya Guo View PDF HTML (experimental) Abstract:Mathematical reasoning poses a significant challenge for language models due to its complex and structured nature. In this paper, we introduce DeepSeekMath 7B, which continues pre-training DeepSeek-Coder-Base-v1.5 7B with 120B math-related tokens sourced from Common Crawl, together with natural language and code data. DeepSeekMath 7B has achieved an impressive score of 51.7% on the competition-level MATH benchmark…
saved by
related reading
- DeepSeek-R1arxiv.org
- As Rocks May Think | Eric Jangevjang.com
- Mathematics in the Library of Babel - Daniel Littdaniellitt.com
- Explore | alphaXivalphaxiv.org
- GRPO Trainer · Hugging Facehuggingface.co
- 2310.10631arxiv.org
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Modelsarxiv.org
- Open-R1: a fully open reproduction of DeepSeek-R1huggingface.co
- [2201.11903] Chain of Thought Prompting Elicits Reasoning in Large Language Modelsarxiv.org
- Training Math Reasoning Model with Reinforcement Learning - NVIDIA ADLRresearch.nvidia.com
- the-illusion-of-thinking.pdfml-site.cdn-apple.com
- The Unreasonable Effectiveness of LLMs in Mathematicschrishayduk.com