✳flâneur — a map of the web's best reading
4.4 Value Iteration
incompleteideas.net · 868 words · saved by 1 readers
4.4 Value Iteration
4.4 Value Iteration Next: 4.5 Asynchronous Dynamic Programming Up: 4. Dynamic Programming Previous: 4.3 Policy Iteration Contents 4.4 Value Iteration One drawback to policy iteration is that each of its iterations involves policy evaluation, which may itself be a protracted iterative computation requiring multiple sweeps through the state set. If policy evaluation is done iteratively, then convergence exactly to occurs only in the limit. Must we wait for exact convergence, or can we stop short of that? The example in Figure 4.2 certainly suggests that it may be possible to truncate policy eval
Explore this link on the map →saved by
related reading
- A (Long) Peek into Reinforcement Learning | Lil'Loglilianweng.github.io
- An Updated Introduction to Reinforcement Learning | Sri's Blogsrianumakonda.com
- Policy Gradient Algorithms | Lil'Loglilianweng.github.io
- Part 1: Key Concepts in RL - Spinning Up documentationspinningup.openai.com
- Part 3: Intro to Policy Optimization - Spinning Up documentationspinningup.openai.com
- Vanilla Policy Gradient - Spinning Up documentationspinningup.openai.com
- Reinforcement Learning in Newcomblike Problemsproceedings.neurips.cc
- Why Momentum Really Worksdistill.pub
- Technical Note: Q-Learning | Machine Learning | Springer Nature Linklink.springer.com
- Bellman equation - Wikipediaen.wikipedia.org
- RL without TD learningseohong.me
- pistar06.pdfpi.website