Life lessons from reinforcement learning — Jason Wei
Becoming an RL diehard in the past year and thinking about RL for most of my waking hours inadvertently taught me an important lesson about how to live my own life. One of the big concepts in RL is that you always want to be “on-policy”: instead of mimicking other people’s successful trajectories,
Life lessons from reinforcement learning Jul 15 Written By Jason Wei Becoming an RL diehard in the past year and thinking about RL for most of my waking hours inadvertently taught me an important lesson about how to live my own life. One of the big concepts in RL is that you always want to be “on-policy”: instead of mimicking other people’s successful trajectories, you should take your own actions and learn from the reward given by the environment. Obviously imitation learning is useful to bootstrap to nonzero pass rate initially, but once you can take reasonable trajectories, we generally avo
Explore this link on the map →related reading
- On-Policy Distillation - Thinking Machines Labthinkingmachines.ai
- Pedagogical RL: Teaching Models to Teach Themselves from Privileged Information - Noah Ziemsnoahziems.com
- Should I Use Offline RL or Imitation Learning? – The Berkeley Artificial Intelligence Research Blogbair.berkeley.edu
- Part 1: Key Concepts in RL - Spinning Up documentationspinningup.openai.com
- A (Long) Peek into Reinforcement Learning | Lil'Loglilianweng.github.io
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- Learning to Imitate | SAIL Blogai.stanford.edu
- A VLA that Learns from Experiencepi.website
- Reinforcement Learning via Implicit Imitation Guidancearxiv.org
- State of Robot Learning, December 2025vedder.io
- pistar06.pdfpi.website
- π*0.6: a VLA That Learns From Experiencephysicalintelligence.company