✳flâneur — a map of the web's best reading
The State of Reinforcement Learning for LLM Reasoning
magazine.sebastianraschka.com · 8,009 words · saved by 1 readers
Understanding GRPO and New Insights from Reasoning Model Papers
The State of Reinforcement Learning for LLM Reasoning Understanding GRPO and New Insights from Reasoning Model Papers Sebastian Raschka, PhD Apr 19, 2025 518 35 40 Share A lot has happened this month, especially with the releases of new flagship models like GPT-4.5 and Llama 4. But you might have noticed that reactions to these releases were relatively muted. Why? One reason could be that GPT-4.5 and Llama 4 remain conventional models, which means they were trained without explicit reinforcement learning for reasoning. Meanwhile, competitors such as xAI and Anthropic have added more reasoning
Explore this link on the map →saved by
related reading
- A Guide to Reinforcement Learning Post-Training for LLMs: PPO, DPO, GRPO, and Beyondhuggingface.co
- rlhfbook.com/book.pdfrlhfbook.com
- DeepSeek-R1arxiv.org
- A vision researcher’s guide to some RL stuff: PPO & GRPO - Yuge (Jimmy) Shiyugeten.github.io
- The State of Reinforcement Learning for LLM Reasoningsebastianraschka.com
- State of RL for reasoning LLMs | A. Weersaweers.de
- DeepSeek R1's recipe to replicate o1 and the future of reasoning LMsinterconnects.ai
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- Understanding R1-Zero-Like Training: A Critical Perspectivearxiv.org
- [2501.12948] DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learningarxiv.org
- Understanding Reasoning LLMs - by Sebastian Raschka, PhDsebastianraschka.com
- Explore | alphaXivalphaxiv.org