flâneur — a map of the web's best reading

Scalable agent alignment via reward modeling | by DeepMind Safety Research | Medium

deepmindsafetyresearch.medium.com · 1,396 words · saved by 1 readers

This post provides an overview of our new paper that outlines a research direction for solving the agent alignment problem. Our approach relies on the recursive application of reward modeling to solve complex real-world problems in a way that aligns with user intentions. In recent years, reinforcement learning has yielded impressive performance in complex game environments ranging from Atari, Go, and chess to Dota 2 and StarCraft II, with artificial agents rapidly surpassing the human level of play in increasingly complex domains. Games are an ideal platform for developing and testing machine learning algorithms. They present challenging tasks that require a range of cognitive abilities to accomplish, mirroring skills needed to solve problems in the real world. Machine learning researchers can run thousands of simulated experiments on the cloud in parallel, generating as much training data as needed for the system to learn. Crucially, games often have a clear objective, and a score tha

Machine Learning Artificial Intelligence Alignment Reinforcement Learning Scalable agent alignment via reward modeling DeepMind Safety Research 6 min read · Nov 20, 2018 -- 3 Listen Share By Jan Leike This post provides an overview of our new paper that outlines a research direction for solving the agent alignment problem. Our approach relies on the recursive application of reward modeling to solve complex real-world problems in a way that aligns with user intentions. In recent years, reinforcement learning has yielded impressive performance in complex game environments ranging from Atari , Go

Explore this link on the map →

related reading