[1606.03476] Generative Adversarial Imitation Learning
Consider learning a policy from example expert behavior, without interaction with the expert or access to reinforcement signal. One approach is to recover the expert's cost function with inverse reinforcement learning, then extract a policy from that cost function with reinforcement learning. This approach is indirect and can be slow. We propose a new general framework for directly extracting a policy from data, as if it were obtained by reinforcement learning following inverse reinforcement learning. We show that a certain instantiation of our framework draws an analogy between imitation learning and generative adversarial networks, from which we derive a model-free imitation learning algorithm that obtains significant performance gains over existing model-free methods in imitating complex behaviors in large, high-dimensional environments.
View PDF HTML (experimental) Abstract:Consider learning a policy from example expert behavior, without interaction with the expert or access to reinforcement signal. One approach is to recover the expert's cost function with inverse reinforcement learning, then extract a policy from that cost function with reinforcement learning. This approach is indirect and can be slow. We propose a new general framework for directly extracting a policy from data, as if it were obtained by reinforcement learning following inverse reinforcement learning. We show that a certain instantiation of our…
saved by
related reading
- InfoGAIL: Interpretable Imitation Learning from Visual Demonstrationsarxiv.org
- Generative Adversarial Imitation Learning.pdfcs.stanford.edu
- Learning to Imitate | SAIL Blogai.stanford.edu
- Ch. 21 - Imitation Learningunderactuated.mit.edu
- Imitating Latent Policies from Observationarxiv.org
- Just Ask for Generalization | Eric Jangevjang.com
- State of Robot Learning, December 2025vedder.io
- Reinforcement Learning via Implicit Imitation Guidancearxiv.org
- 2312.10812arxiv.org
- 1011.0686arxiv.org
- Life lessons from reinforcement learning - Jason Weijasonwei.net
- A (Long) Peek into Reinforcement Learning | Lil'Loglilianweng.github.io