The State of Reinforcement Learning for LLM Reasoning
A lot has happened this month, especially with the releases of new flagship models like GPT-4.5 and Llama 4. But you might have noticed that reactions to these releases were relatively muted. Why? One reason could be that GPT-4.5 and Llama 4 remain conventional models, which means they were trained without explicit reinforcement learning for reasoning. Meanwhile, competitors such as xAI and Anthropic have added more reasoning capabilities and features into their models. For instance, both the xAI Grok and Anthropic Claude interfaces now include a “thinking” (or “extended thinking”) button for certain models that explicitly toggles reasoning capabilities. In any case, the muted response to GPT-4.5 and Llama 4 (non-reasoning) models suggests we are approaching the limits of what scaling model size and data alone can achieve. However, OpenAI’s recent release of the o3 reasoning model demonstrates there is still considerable room for improvement when investing compute strategically, specif
The State of Reinforcement Learning for LLM Reasoning Understanding GRPO and New Insights from Reasoning Model Papers Sebastian Raschka, PhD Apr 19, 2025 524 36 40 Share A lot has happened this month, especially with the releases of new flagship models like GPT-4.5 and Llama 4. But you might have noticed that reactions to these releases were relatively muted. Why? One reason could be that GPT-4.5 and Llama 4 remain conventional models, which means they were trained without explicit reinforcement learning for reasoning. Meanwhile, competitors such as xAI and Anthropic have added more reasoning
related reading
- The State of Reinforcement Learning for LLM Reasoningmagazine.sebastianraschka.com
- DeepSeek-R1arxiv.org
- State of RL for reasoning LLMs | A. Weersaweers.de
- [2504.13837] Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?arxiv.org
- As Rocks May Think | Eric Jangevjang.com
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- DeepSeek R1's recipe to replicate o1 and the future of reasoning LMsinterconnects.ai
- GRPO++: Tricks for Making RL Actually Workcameronrwolfe.substack.com
- [2501.12948] DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learningarxiv.org
- A vision researcher’s guide to some RL stuff: PPO & GRPO - Yuge (Jimmy) Shiyugeten.github.io
- Understanding Reasoning LLMs - by Sebastian Raschka, PhDsebastianraschka.com
- Understanding R1-Zero-Like Training: A Critical Perspectivearxiv.org