Proximal Policy Optimization | OpenAI
We use cookies and similar technologies to deliver, maintain, improve our services and for security purposes. Check our Privacy Policy for details. Click 'Accept all' to let OpenAI and partners use cookies for these purposes. Click 'Reject non-essential' to say no to cookies, except those that are strictly necessary. July 20, 2017 Illustration: Ben Barry We’re releasing a new class of reinforcement learning algorithms, Proximal Policy Optimization (PPO), which perform comparably or better than state-of-the-art approaches while being much simpler to implement and tune. PPO has become the default reinforcement learning algorithm at OpenAI because of its ease of use and good performance. Policy gradient methods (opens in a new window) are fundamental to recent breakthroughs in using deep neural networks for control, from video games (opens in a new window) , to 3D locomotion (opens in a new window) , to Go (opens in a new window) . But getting good results via policy gradient methods is
We use cookies and similar technologies to deliver, maintain, improve our services and for security purposes. Check our Privacy Policy for details. Click 'Accept all' to let OpenAI and partners use cookies for these purposes. Click 'Reject non-essential' to say no to cookies, except those that are strictly necessary. July 20, 2017 Illustration: Ben Barry We’re releasing a new class of reinforcement learning algorithms, Proximal Policy Optimization (PPO), which perform comparably or better than state-of-the-art approaches while being much simpler to implement and tune. PPO has become the defaul
Explore this link on the map →