Reinforcement Learning via Implicit Imitation Guidance
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions. HTML conversions sometimes display errors due to content that did not convert correctly from the source. This paper uses the following packages that are not yet supported by the HTML conversion tool. Feedback on these issues are not necessary; they are known and are being worked on. Authors: achieve the best HTML results from your LaTeX submissions by following these best practices. We study the problem of sample efficient reinforcement learning, where prior data such as demonstrations are provided for initialization in lieu of a dense reward signal. A natural approach is to incorporate an imitation learning objective, either as regularization during training or to acquire a reference policy. However, imitation learning objectives can ultimately degrade
Reinforcement Learning via Implicit Imitation Guidance Perry Dong*, Alec M. Lessing*, Annie S. Chen*, Chelsea Finn Stanford University; {perryd, aleclessing, asc8}@stanford.edu Abstract We study the problem of sample efficient reinforcement learning, where prior data such as demonstrations are provided for initialization in lieu of a dense reward signal. A natural approach is to incorporate an imitation learning objective, either as regularization during training or to acquire a reference policy. However, imitation learning objectives can ultimately degrade long-term performance, as it does no
Explore this link on the map →saved by
related reading
- State of Robot Learning, December 2025vedder.io
- Should I Use Offline RL or Imitation Learning? – The Berkeley Artificial Intelligence Research Blogbair.berkeley.edu
- Ch. 21 - Imitation Learningunderactuated.mit.edu
- Imitation Bootstrapped Reinforcement Learningarxiv.org
- Learning to Imitate | SAIL Blogai.stanford.edu
- Exploration for the Efficient Deployment of Reinforcement Learning Agentsopenreview.net
- State of RL for reasoning LLMs | A. Weersaweers.de
- A vision researcher’s guide to some RL stuff: PPO & GRPO - Yuge (Jimmy) Shiyugeten.github.io
- Key Papers in Deep RL - Spinning Up documentationspinningup.openai.com
- pistar06.pdfpi.website
- Deep Q-Networks Explained — LessWronglesswrong.com
- Exploration Strategies in Deep Reinforcement Learning | Lil'Loglilianweng.github.io