flâneur — a map of the web's best reading

RUDDER - Reinforcement Learning with Delayed Rewards | rudder

ml-jku.github.io · 4,434 words · saved by 1 readers

Blog post

RUDDER - Reinforcement Learning with Delayed Rewards | rudder Blog post to RUDDER: Return Decomposition for Delayed Rewards . Recently, tasks with delayed rewards that required model-free reinforcement learning attracted a lot of attention via complex strategy games. For example, DeepMind currently focuses on the delayed reward games Capture the flag and Starcraft , whereas Microsoft is putting up the Marlo environment, and Open AI announced its Dota 2 achievements. Mastering these games with delayed rewards using model-free reinforcement learning poses a great challenge and an almost insurmou

Explore this link on the map →

saved by

related reading