Scalable agent alignment via reward modeling | by DeepMind Safety Research | Medium
This post provides an overview of our new paper that outlines a research direction for solving the agent alignment problem. Our approach relies on the recursive application of reward modeling to solve complex real-world problems in a way that aligns with user intentions. In recent years, reinforcement learning has yielded impressive performance in complex game environments ranging from Atari, Go, and chess to Dota 2 and StarCraft II, with artificial agents rapidly surpassing the human level of play in increasingly complex domains. Games are an ideal platform for developing and testing machine learning algorithms. They present challenging tasks that require a range of cognitive abilities to accomplish, mirroring skills needed to solve problems in the real world. Machine learning researchers can run thousands of simulated experiments on the cloud in parallel, generating as much training data as needed for the system to learn. Crucially, games often have a clear objective, and a score tha
Machine Learning Artificial Intelligence Alignment Reinforcement Learning Scalable agent alignment via reward modeling DeepMind Safety Research 6 min read · Nov 20, 2018 -- 3 Listen Share By Jan Leike This post provides an overview of our new paper that outlines a research direction for solving the agent alignment problem. Our approach relies on the recursive application of reward modeling to solve complex real-world problems in a way that aligns with user intentions. In recent years, reinforcement learning has yielded impressive performance in complex game environments ranging from Atari , Go
Explore this link on the map →related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Research Areas in Evaluation and Guarantees in Reinforcement Learning (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Research Areas in Methods for Post-training and Elicitation (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- Why I’m optimistic about our alignment approachaligned.substack.com
- Reward is not the optimization target — LessWronglesswrong.com
- Teaching Claude why \ Anthropicanthropic.com
- Models Don't "Get Reward" — LessWronglesswrong.com
- Why I’m optimistic about our alignment approachaligned.substack.com
- [AN #70]: Agents that help humans who are still learning about their own preferences — LessWronglesswrong.com
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com