flâneur — a map of the web's best reading

The State of Reinforcement Learning for LLM Reasoning

magazine.sebastianraschka.com · 8,009 words · saved by 1 readers

Understanding GRPO and New Insights from Reasoning Model Papers

The State of Reinforcement Learning for LLM Reasoning Understanding GRPO and New Insights from Reasoning Model Papers Sebastian Raschka, PhD Apr 19, 2025 518 35 40 Share A lot has happened this month, especially with the releases of new flagship models like GPT-4.5 and Llama 4. But you might have noticed that reactions to these releases were relatively muted. Why? One reason could be that GPT-4.5 and Llama 4 remain conventional models, which means they were trained without explicit reinforcement learning for reasoning. Meanwhile, competitors such as xAI and Anthropic have added more reasoning

Explore this link on the map →

saved by

related reading