flâneur — a map of the web's best reading

ar0cket1 on X: "Solving OPSD (basically)" / X

x.com · 667 words · saved by 2 readers

https://t.co/QL4SkUecDr

@ar0cket1: Solving OPSD (basically) self hinted teachers will likely be common practice in a few months for RL, here is the progress and findings I have made: *note: this is a continuation of my initial blog ar0cket1 @ar0cket1 · May 12 Article On Policy Self Distillation I’ve been working on solving OPSD for a week now (and a little more). The goal is RL like upper bound, with OPD like sample efficiency while being stable. This is the results I’ve gotten so far +... 6 9 142 53K Context Just to contextualize everything: OPSD’s goal can be thought of as achieving OPD like sample efficiency w

Explore this link on the map →

saved by

related reading