AlphaGo Zero: Minimal Policy Improvement, Expectation Propagation and other Connections
This is a post about the new reinforcement learning technique that enables AlphaGo Zero to learn Go from scratch via self-play. The paper has been out for a week I guess it's now considered old - sorry for the latency. D Silver, J Schrittwieser, K Simonyan, I Antonoglou, A Huang,
October 26, 2017 AlphaGo Zero: Minimal Policy Improvement, Expectation Propagation and other Connections This is a post about the new reinforcement learning technique that enables AlphaGo Zero to learn Go from scratch via self-play. The paper has been out for a week I guess it's now considered old - sorry for the latency. D Silver, J Schrittwieser, K Simonyan, I Antonoglou, A Huang, Arthur Guez, T Hubert, L Baker, M Lai, Adrian Bolton, Y Chen, T Lillicrap, F Hui, L Sifre, G van den Driessche, T Graepel & D Hassabis (2017) Mastering the Game of Go without Human Knowledge I'm no expert in RL, so
Explore this link on the map →saved by
related reading
- How DeepMind's Generally Capable Agents Were Trained — LessWronglesswrong.com
- A (Long) Peek into Reinforcement Learning | Lil'Loglilianweng.github.io
- Simple Alpha Zeroweb.stanford.edu
- EfficientZero: How It Works — LessWronglesswrong.com
- A vision researcher’s guide to some RL stuff: PPO & GRPO - Yuge (Jimmy) Shiyugeten.github.io
- Policy Gradient Algorithms | Lil'Loglilianweng.github.io
- The 37 Implementation Details of Proximal Policy Optimization · The ICLR Blog Trackiclr-blog-track.github.io
- Part 3: Intro to Policy Optimization - Spinning Up documentationspinningup.openai.com
- State of RL for reasoning LLMs | A. Weersaweers.de
- An Updated Introduction to Reinforcement Learning | Sri's Blogsrianumakonda.com
- [1707.06347] Proximal Policy Optimization Algorithmsarxiv.org
- RLHF Bookrlhfbook.com