Pedagogical RL: Teaching Models to Teach Themselves from Privileged Information - Noah Ziems
Souradip Chakraborty*,1,2, Noah Ziems*,1,3, Furong Huang2, Meng Jiang3, Amrit Singh Bedi4, Omar Khattab1 1MIT 2UMD 3UND 4UCF *Equal contribution
Souradip Chakraborty *,1,2 , Noah Ziems *,1,3 , Furong Huang 2 , Meng Jiang 3 , Amrit Singh Bedi 4 , Omar Khattab 1 1 MIT 2 UMD 3 ND 4 UCF * Equal contribution May 14, 2026 · Announcement tweet Typical reinforcement learning and on-policy distillation algorithms rely on privileged information like labeled final answers or execution feedback to evaluate rollouts , but do not actually benefit from them for finding good rollouts . If your model can’t already stumble upon successful trajectories, RL simply stalls. In this post, we ask: Can we leverage privileged information to actively samp
saved by
related reading
- On-Policy Distillation - Thinking Machines Labthinkingmachines.ai
- State of RL for reasoning LLMs | A. Weersaweers.de
- Reinforcement Learning for Knowledge Awareness – kalomaze's kalomazing blogkalomaze.bearblog.dev
- will brown on X: "On SFT, RL, and on-policy distillation" / Xx.com
- RLHF Bookrlhfbook.com
- [2604.13010] Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillationarxiv.org
- GRPO++: Tricks for Making RL Actually Workcameronrwolfe.substack.com
- A vision researcher’s guide to some RL stuff: PPO & GRPO - Yuge (Jimmy) Shiyugeten.github.io
- SFT, RL, and On-Policy Distillation Through a Distributional Lens | whnrehiew.github.io
- How can LLM RL Work Despite Information-Theoretic Inefficiencyberen.io
- GRPO is terrible — LessWronglesswrong.com
- Life lessons from reinforcement learning - Jason Weijasonwei.net