GRPO Trainer
huggingface.co · 7,914 words · saved by 1 readers
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
Overview TRL supports the GRPO Trainer for training language models, as described in the paper DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models by Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Mingchuan Zhang, Y. K. Li, Y. Wu, Daya Guo. The abstract from the paper is the following: Mathematical reasoning poses a significant challenge for language models due to its complex and structured nature. In this paper, we introduce DeepSeekMath 7B, which continues pre-training DeepSeek-Coder-Base-v1.5 7B with 120B math-related tokens sourced from…
related reading
- DeepSeek-R1arxiv.org
- GRPO++: Tricks for Making RL Actually Workcameronrwolfe.substack.com
- State of RL for reasoning LLMs | A. Weersaweers.de
- Tutorial: Train your own Reasoning model with GRPO | Unsloth Documentationdocs.unsloth.ai
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Modelsarxiv.org
- Interactive Visualization of RL Algorithms for LLM Trainingzcy233035.github.io
- Why GRPO is Important and How it Worksghost.oxen.ai
- Composer2.pdfcursor.com
- As Rocks May Think | Eric Jangevjang.com
- A vision researcher’s guide to some RL stuff: PPO & GRPO - Yuge (Jimmy) Shiyugeten.github.io
- The State of Reinforcement Learning for LLM Reasoningmagazine.sebastianraschka.com
- Understanding R1-Zero-Like Training: A Critical Perspectivearxiv.org