How to Explore to Scale RL Training of LLMs on Hard Problems? – Machine Learning Blog | ML@CMU | Carnegie Mellon University
LLM RL typically operates in one of three exploration regimes: sharpening, chaining, or guided exploration; standard RL stays in the first two and plateaus on hard problems, even those in the training set. Mixing easy and hard data triggers interference, causing reward-rich tasks to drown out hard o
How to Explore to Scale RL Training of LLMs on Hard Problems? – Machine Learning Blog | ML@CMU | Carnegie Mellon University --> Input your search keywords and press Enter. Categories: Research Educational Categories: Research Educational machine learning reinforcement learning Research How to Explore to Scale RL Training of LLMs on Hard Problems? Authors Amrith Setlur by Amrith Setlur --> Affiliations Published November 26, 2025 DOI Figure 1. Three regimes of exploration: Current RL model can explore via: (1) sharpening : simply increases likelihood on traces it can sample with high prob
Explore this link on the map →saved by
related reading
- IsoCompute Playbook: Optimally Scaling Sampling Compute for RL Training of LLMscompute-optimal-rl-llm-scaling.github.io
- State of RL for reasoning LLMs | A. Weersaweers.de
- DeepSeek-R1arxiv.org
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- A vision researcher’s guide to some RL stuff: PPO & GRPO - Yuge (Jimmy) Shiyugeten.github.io
- The State of Reinforcement Learning for LLM Reasoningmagazine.sebastianraschka.com
- RL Scaling Laws for LLMs - by Cameron R. Wolfe, Ph.D.cameronrwolfe.substack.com
- PPO for LLMs: A Guide for Normal Peoplecameronrwolfe.substack.com
- [2606.10346] Reasoning or Memorization? Direction-Aware Diversity Exploration in LLM Reinforcement Learningarxiv.org
- Optimizing LLM Test-Time Compute Involves Solving a Meta-RL Problem – Machine Learning Blog | ML@CMU | Carnegie Mellon Universityblog.ml.cmu.edu
- From REINFORCE to Dr. GRPOlancelqf.github.io
- [2507.19457] GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learningarxiv.org