flâneur — a map of the web's best reading

Bite: How Deepseek R1 was trained

philschmid.de · 724 words · saved by 1 readers

5 Minute Read on how Deepseek R1 was trained using Group Relative Policy Optimization (GRPO) and RL-focused multi-stage training approach.

Bite: How Deepseek R1 was trained January 17, 2025 4 minute read DeepSeek AI released DeepSeek-R1, an open model that rivals OpenAI's o1 in complex reasoning tasks, introduced using Group Relative Policy Optimization (GRPO) and RL-focused multi-stage training approach. Understanding Group Relative Policy Optimization (GRPO) Group Relative Policy Optimization (GRPO) is a reinforcement learning algorithm to improve the reasoning capabilities of LLMs. It was introduced in the DeepSeekMath paper in the context of mathematical reasoning. GRPO modifies the traditional Proximal Policy Optimization (P

Explore this link on the map →

related reading