The State of Reinforcement Learning for LLM Reasoning
magazine.sebastianraschka.com · 8,009 words · saved by 1 readers
Understanding GRPO and New Insights from Reasoning Model Papers
The State of Reinforcement Learning for LLM Reasoning Understanding GRPO and New Insights from Reasoning Model Papers Sebastian Raschka, PhD Apr 19, 2025 518 35 40 Share A lot has happened this month, especially with the releases of new flagship models like GPT-4.5 and Llama 4. But you might have noticed that reactions to these releases were relatively muted. Why? One reason could be that GPT-4.5 and Llama 4 remain conventional models, which means they were trained without explicit reinforcement learning for reasoning. Meanwhile, competitors such as xAI and Anthropic have added more reasoning
saved by
related reading
- A Guide to Reinforcement Learning Post-Training for LLMs: PPO, DPO, GRPO, and Beyondhuggingface.co
- rlhfbook.com/book.pdfrlhfbook.com
- DeepSeek-R1arxiv.org
- A vision researcher’s guide to some RL stuff: PPO & GRPO - Yuge (Jimmy) Shiyugeten.github.io
- The State of Reinforcement Learning for LLM Reasoningsebastianraschka.com
- State of RL for reasoning LLMs | A. Weersaweers.de
- As Rocks May Think | Eric Jangevjang.com
- DeepSeek R1's recipe to replicate o1 and the future of reasoning LMsinterconnects.ai
- GRPO++: Tricks for Making RL Actually Workcameronrwolfe.substack.com
- [2504.13837] Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?arxiv.org
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- Understanding R1-Zero-Like Training: A Critical Perspectivearxiv.org