Imitation Bootstrapped Reinforcement Learning
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions. HTML conversions sometimes display errors due to content that did not convert correctly from the source. This paper uses the following packages that are not yet supported by the HTML conversion tool. Feedback on these issues are not necessary; they are known and are being worked on. Authors: achieve the best HTML results from your LaTeX submissions by following these best practices. Despite the considerable potential of reinforcement learning (RL), robotic control tasks predominantly rely on imitation learning (IL) due to its better sample efficiency. However, it is costly to collect comprehensive expert demonstrations that enable IL to generalize to all possible scenarios, and any distribution shift would require recollecting data for finetuning. Theref
Imitation Bootstrapped Reinforcement Learning Hengyuan Hu Stanford University Suvir Mirchandani Stanford Univeristy Dorsa Sadigh Stanford University Abstract Despite the considerable potential of reinforcement learning (RL), robotic control tasks predominantly rely on imitation learning (IL) due to its better sample efficiency. However, it is costly to collect comprehensive expert demonstrations that enable IL to generalize to all possible scenarios, and any distribution shift would require recollecting data for finetuning. Therefore, RL is appealing if it can build upon IL as an efficient aut
Explore this link on the map →related reading
- Reinforcement Learning via Implicit Imitation Guidancearxiv.org
- State of Robot Learning, December 2025vedder.io
- Precise Manipulation with Efficient Online RLpi.website
- pistar06.pdfpi.website
- Should I Use Offline RL or Imitation Learning? – The Berkeley Artificial Intelligence Research Blogbair.berkeley.edu
- Ch. 21 - Imitation Learningunderactuated.mit.edu
- A (Long) Peek into Reinforcement Learning | Lil'Loglilianweng.github.io
- Part 1: Key Concepts in RL - Spinning Up documentationspinningup.openai.com
- State of RL for reasoning LLMs | A. Weersaweers.de
- The 37 Implementation Details of Proximal Policy Optimization · The ICLR Blog Trackiclr-blog-track.github.io
- Learning to Imitate | SAIL Blogai.stanford.edu
- π*0.6: a VLA That Learns From Experiencephysicalintelligence.company