The State of Reinforcement Learning for LLM Reasoning
A lot has happened this month, especially with the releases of new flagship models like GPT-4.5 and Llama 4. But you might have noticed that reactions to these releases were relatively muted. Why? One reason could be that GPT-4.5 and Llama 4 remain conventional models, which means they were trained without explicit reinforcement learning for reasoning. Meanwhile, competitors such as xAI and Anthropic have added more reasoning capabilities and features into their models. For instance, both the xAI Grok and Anthropic Claude interfaces now include a “thinking” (or “extended thinking”) button for certain models that explicitly toggles reasoning capabilities. In any case, the muted response to GPT-4.5 and Llama 4 (non-reasoning) models suggests we are approaching the limits of what scaling model size and data alone can achieve. However, OpenAI’s recent release of the o3 reasoning model demonstrates there is still considerable room for improvement when investing compute strategically, specif
The State of Reinforcement Learning for LLM Reasoning Understanding GRPO and New Insights from Reasoning Model Papers Sebastian Raschka, PhD Apr 19, 2025 524 36 40 Share A lot has happened this month, especially with the releases of new flagship models like GPT-4.5 and Llama 4. But you might have noticed that reactions to these releases were relatively muted. Why? One reason could be that GPT-4.5 and Llama 4 remain conventional models, which means they were trained without explicit reinforcement learning for reasoning. Meanwhile, competitors such as xAI and Anthropic have added more reasoning
Explore this link on the map →related reading
- The State of Reinforcement Learning for LLM Reasoningmagazine.sebastianraschka.com
- DeepSeek-R1arxiv.org
- State of RL for reasoning LLMs | A. Weersaweers.de
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- DeepSeek R1's recipe to replicate o1 and the future of reasoning LMsinterconnects.ai
- A vision researcher’s guide to some RL stuff: PPO & GRPO - Yuge (Jimmy) Shiyugeten.github.io
- [2501.12948] DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learningarxiv.org
- Understanding Reasoning LLMs - by Sebastian Raschka, PhDsebastianraschka.com
- Understanding R1-Zero-Like Training: A Critical Perspectivearxiv.org
- Explore | alphaXivalphaxiv.org
- [2606.10346] Reasoning or Memorization? Direction-Aware Diversity Exploration in LLM Reinforcement Learningarxiv.org
- LLM Resourcesforrestbicker.com