Pedagogical RL: Teaching Models to Teach Themselves from Privileged Information - Noah Ziems
Souradip Chakraborty*,1,2, Noah Ziems*,1,3, Furong Huang2, Meng Jiang3, Amrit Singh Bedi4, Omar Khattab1 1MIT 2UMD 3UND 4UCF *Equal contribution
Souradip Chakraborty *,1,2 , Noah Ziems *,1,3 , Furong Huang 2 , Meng Jiang 3 , Amrit Singh Bedi 4 , Omar Khattab 1 1 MIT 2 UMD 3 ND 4 UCF * Equal contribution May 14, 2026 · Announcement tweet Typical reinforcement learning and on-policy distillation algorithms rely on privileged information like labeled final answers or execution feedback to evaluate rollouts , but do not actually benefit from them for finding good rollouts . If your model can’t already stumble upon successful trajectories, RL simply stalls. In this post, we ask: Can we leverage privileged information to actively samp
Explore this link on the map →saved by
related reading
- On-Policy Distillation - Thinking Machines Labthinkingmachines.ai
- State of RL for reasoning LLMs | A. Weersaweers.de
- Reinforcement Learning for Knowledge Awareness – kalomaze's kalomazing blogkalomaze.bearblog.dev
- RLHF Bookrlhfbook.com
- [2604.13010] Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillationarxiv.org
- A vision researcher’s guide to some RL stuff: PPO & GRPO - Yuge (Jimmy) Shiyugeten.github.io
- GRPO is terrible — LessWronglesswrong.com
- SFT, RL, and On-Policy Distillation Through a Distributional Lens | whnrehiew.github.io
- [2605.23857] Strong Teacher Not Needed? On Distillation in LLM Pretrainingarxiv.org
- [2601.18734] Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Modelsarxiv.org
- Learning Beyond Gradientstrinkle23897.github.io
- [2605.10889] Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Whyarxiv.org