Bite: How Deepseek R1 was trained
philschmid.de · 724 words · saved by 1 readers
5 Minute Read on how Deepseek R1 was trained using Group Relative Policy Optimization (GRPO) and RL-focused multi-stage training approach.
Bite: How Deepseek R1 was trained January 17, 2025 4 minute read DeepSeek AI released DeepSeek-R1, an open model that rivals OpenAI's o1 in complex reasoning tasks, introduced using Group Relative Policy Optimization (GRPO) and RL-focused multi-stage training approach. Understanding Group Relative Policy Optimization (GRPO) Group Relative Policy Optimization (GRPO) is a reinforcement learning algorithm to improve the reasoning capabilities of LLMs. It was introduced in the DeepSeekMath paper in the context of mathematical reasoning. GRPO modifies the traditional Proximal Policy Optimization (P
related reading
- Why GRPO is Important and How it Worksghost.oxen.ai
- The Math Behind DeepSeek: A Deep Dive into Group Relative Policy Optimization (GRPO)medium.com
- DeepSeek-R1arxiv.org
- GRPO++: Tricks for Making RL Actually Workcameronrwolfe.substack.com
- A vision researcher’s guide to some RL stuff: PPO & GRPO - Yuge (Jimmy) Shiyugeten.github.io
- State of RL for reasoning LLMs | A. Weersaweers.de
- Interactive Visualization of RL Algorithms for LLM Trainingzcy233035.github.io
- DeepSeek R1's recipe to replicate o1 and the future of reasoning LMsinterconnects.ai
- GRPO Trainer · Hugging Facehuggingface.co
- The State of Reinforcement Learning for LLM Reasoningmagazine.sebastianraschka.com
- Tutorial: Train your own Reasoning model with GRPO | Unsloth Documentationdocs.unsloth.ai
- Open-R1: a fully open reproduction of DeepSeek-R1huggingface.co