flâneur — a map of the web's best reading

AlphaGo Zero: Minimal Policy Improvement, Expectation Propagation and other Connections

inference.vc · 1,670 words · saved by 1 readers

This is a post about the new reinforcement learning technique that enables AlphaGo Zero to learn Go from scratch via self-play. The paper has been out for a week I guess it's now considered old - sorry for the latency. D Silver, J Schrittwieser, K Simonyan, I Antonoglou, A Huang,

October 26, 2017 AlphaGo Zero: Minimal Policy Improvement, Expectation Propagation and other Connections This is a post about the new reinforcement learning technique that enables AlphaGo Zero to learn Go from scratch via self-play. The paper has been out for a week I guess it's now considered old - sorry for the latency. D Silver, J Schrittwieser, K Simonyan, I Antonoglou, A Huang, Arthur Guez, T Hubert, L Baker, M Lai, Adrian Bolton, Y Chen, T Lillicrap, F Hui, L Sifre, G van den Driessche, T Graepel & D Hassabis (2017) Mastering the Game of Go without Human Knowledge I'm no expert in RL, so

Explore this link on the map →

saved by

related reading