How to Explore to Scale RL Training of LLMs on Hard Problems? – Machine Learning Blog | ML@CMU | Carnegie Mellon University
LLM RL typically operates in one of three exploration regimes: sharpening, chaining, or guided exploration; standard RL stays in the first two and plateaus on hard problems, even those in the training set. Mixing easy and hard data triggers interference, causing reward-rich tasks to drown out hard o
How to Explore to Scale RL Training of LLMs on Hard Problems? – Machine Learning Blog | ML@CMU | Carnegie Mellon University --> Input your search keywords and press Enter. Categories: Research Educational Categories: Research Educational machine learning reinforcement learning Research How to Explore to Scale RL Training of LLMs on Hard Problems? Authors Amrith Setlur by Amrith Setlur --> Affiliations Published November 26, 2025 DOI Figure 1. Three regimes of exploration: Current RL model can explore via: (1) sharpening : simply increases likelihood on traces it can sample with high prob
saved by
related reading
- IsoCompute Playbook: Optimally Scaling Sampling Compute for RL Training of LLMscompute-optimal-rl-llm-scaling.github.io
- State of RL for reasoning LLMs | A. Weersaweers.de
- How can LLM RL Work Despite Information-Theoretic Inefficiencyberen.io
- DeepSeek-R1arxiv.org
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- GRPO++: Tricks for Making RL Actually Workcameronrwolfe.substack.com
- A vision researcher’s guide to some RL stuff: PPO & GRPO - Yuge (Jimmy) Shiyugeten.github.io
- As Rocks May Think | Eric Jangevjang.com
- [2602.19362] LLMs Can Learn to Reason Via Off-Policy RLarxiv.org
- The State of Reinforcement Learning for LLM Reasoningmagazine.sebastianraschka.com
- RL Scaling Laws for LLMs - by Cameron R. Wolfe, Ph.D.cameronrwolfe.substack.com
- PPO for LLMs: A Guide for Normal Peoplecameronrwolfe.substack.com