Temporal-Difference Networks
We introduce a generalization of temporal-difference (TD) learning to networks of interrelated predictions. Rather than relating a single pre- diction to itself at a later time, as in conventional TD methods, a TD network relates each prediction in a set of predictions to other predic- tions in the set at a later time. TD networks can represent and apply TD learning to a much wider class of predictions than has previously been possible. Using a random-walk example, we show that these networks can be used to learn to predict by a fixed interval, which is not possi- ble with conventional TD methods. Secondly, we show that if the inter- predictive relationships are made conditional on action, then the usual learning-efficiency advantage of TD methods over Monte Carlo (super- vised learning) methods becomes particularly pronounced. Thirdly, we demonstrate that TD networks can learn predictive state representations that enable exact solution of a non-Markov problem. A very broad range of in
Temporal-Difference Networks Bibtex Metadata Paper Abstract We introduce a generalization of temporal-difference (TD) learning to networks of interrelated predictions. Rather than relating a single pre- diction to itself at a later time, as in conventional TD methods, a TD network relates each prediction in a set of predictions to other predic- tions in the set at a later time. TD networks can represent and apply TD learning to a much wider class of predictions than has previously been possible. Using a random-walk example, we show that these networks can be used to learn to predict by a fixed
Explore this link on the map →related reading
- A (Long) Peek into Reinforcement Learning | Lil'Loglilianweng.github.io
- Deep Q-Networks Explained — LessWronglesswrong.com
- pdfopenreview.net
- RUDDER - Reinforcement Learning with Delayed Rewards | rudderml-jku.github.io
- [1802.09081] Temporal Difference Models: Model-Free Deep RL for Model-Based Controlar5iv.labs.arxiv.org
- An Updated Introduction to Reinforcement Learning | Sri's Blogsrianumakonda.com
- Continuous Thought Machinespub.sakana.ai
- Introducing Dreamer: Scalable Reinforcement Learning Using World Modelsresearch.google
- The Promise of Hierarchical Reinforcement Learningthegradient.pub
- Building machines that learn and think like people | Behavioral and Brain Sciences | Cambridge Corecambridge.org
- [1606.05312] Successor Features for Transfer in Reinforcement Learningarxiv.org
- RL without TD learningseohong.me