flâneur — a map of the web's best reading

How to Explore to Scale RL Training of LLMs on Hard Problems? – Machine Learning Blog | ML@CMU | Carnegie Mellon University

blog.ml.cmu.edu · 6,328 words · saved by 1 readers

LLM RL typically operates in one of three exploration regimes: sharpening, chaining, or guided exploration; standard RL stays in the first two and plateaus on hard problems, even those in the training set. Mixing easy and hard data triggers interference, causing reward-rich tasks to drown out hard o

How to Explore to Scale RL Training of LLMs on Hard Problems? – Machine Learning Blog | ML@CMU | Carnegie Mellon University --> Input your search keywords and press Enter. Categories: Research Educational Categories: Research Educational machine learning reinforcement learning Research How to Explore to Scale RL Training of LLMs on Hard Problems? Authors Amrith Setlur by Amrith Setlur --> Affiliations Published November 26, 2025 DOI Figure 1. Three regimes of exploration: Current RL model can explore via: (1) sharpening : simply increases likelihood on traces it can sample with high prob

Explore this link on the map →

saved by

related reading