✳flâneur — a map of the web's best reading
Bite: How Deepseek R1 was trained
philschmid.de · 724 words · saved by 1 readers
5 Minute Read on how Deepseek R1 was trained using Group Relative Policy Optimization (GRPO) and RL-focused multi-stage training approach.
Bite: How Deepseek R1 was trained January 17, 2025 4 minute read DeepSeek AI released DeepSeek-R1, an open model that rivals OpenAI's o1 in complex reasoning tasks, introduced using Group Relative Policy Optimization (GRPO) and RL-focused multi-stage training approach. Understanding Group Relative Policy Optimization (GRPO) Group Relative Policy Optimization (GRPO) is a reinforcement learning algorithm to improve the reasoning capabilities of LLMs. It was introduced in the DeepSeekMath paper in the context of mathematical reasoning. GRPO modifies the traditional Proximal Policy Optimization (P
Explore this link on the map →related reading
- Why GRPO is Important and How it Worksghost.oxen.ai
- DeepSeek-R1arxiv.org
- A vision researcher’s guide to some RL stuff: PPO & GRPO - Yuge (Jimmy) Shiyugeten.github.io
- State of RL for reasoning LLMs | A. Weersaweers.de
- DeepSeek R1's recipe to replicate o1 and the future of reasoning LMsinterconnects.ai
- The State of Reinforcement Learning for LLM Reasoningmagazine.sebastianraschka.com
- Tutorial: Train your own Reasoning model with GRPO | Unsloth Documentationdocs.unsloth.ai
- Understanding R1-Zero-Like Training: A Critical Perspectivearxiv.org
- Lightweight Guide to understanding GRPO and RL principles - Musings of Muraligitlostmurali.com
- [2605.03327] DGPO: Distribution Guided Policy Optimization for Fine Grained Credit Assignmentarxiv.org
- A Guide to Reinforcement Learning Post-Training for LLMs: PPO, DPO, GRPO, and Beyondhuggingface.co
- GitHub - deepseek-ai/DeepSeek-R1 · GitHubgithub.com