✳flâneur — a map of the web's best reading
ar0cket1 on X: "Solving OPSD (basically)" / X
x.com · 667 words · saved by 2 readers
https://t.co/QL4SkUecDr
@ar0cket1: Solving OPSD (basically) self hinted teachers will likely be common practice in a few months for RL, here is the progress and findings I have made: *note: this is a continuation of my initial blog ar0cket1 @ar0cket1 · May 12 Article On Policy Self Distillation I’ve been working on solving OPSD for a week now (and a little more). The goal is RL like upper bound, with OPD like sample efficiency while being stable. This is the results I’ve gotten so far +... 6 9 142 53K Context Just to contextualize everything: OPSD’s goal can be thought of as achieving OPD like sample efficiency w
Explore this link on the map →saved by
related reading
- SFT, RL, and On-Policy Distillation Through a Distributional Lens | whnrehiew.github.io
- On-Policy Distillation - Thinking Machines Labthinkingmachines.ai
- Pedagogical RL: Teaching Models to Teach Themselves from Privileged Information - Noah Ziemsnoahziems.com
- [2604.13010] Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillationarxiv.org
- [2601.18734] Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Modelsarxiv.org
- Do your capabilities homework — LessWronglesswrong.com
- State of RL for reasoning LLMs | A. Weersaweers.de
- [2604.00626] A Survey of On-Policy Distillation for Large Language Modelsarxiv.org
- Did Claude 3 Opus align itself via gradient hacking? — LessWronglesswrong.com
- [2605.10889] Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Whyarxiv.org
- will brown on X: "On SFT, RL, and on-policy distillation" / Xx.com
- A vision researcher’s guide to some RL stuff: PPO & GRPO - Yuge (Jimmy) Shiyugeten.github.io