[1805.00899] AI safety via debate
To make AI systems broadly useful for challenging real-world tasks, we need them to learn complex human goals and preferences. One approach to specifying complex goals asks humans to judge during training which agent behaviors are safe and useful, but this approach can fail if the task is too complicated for a human to directly judge. To help address this concern, we propose training agents via self play on a zero sum debate game. Given a question or proposed action, two agents take turns making short statements up to a limit, then a human judges which of the agents gave the most true, useful information. In an analogy to complexity theory, debate with optimal play can answer any question in \PSPACE given polynomial time judges (direct judging answers only \NP questions). In practice, whether debate works involves empirical questions about humans and the tasks we want AIs to perform, plus theoretical questions about the meaning of AI alignment. We report results on an initial MNIST
\newclass \DEBATE DEBATE \newclass \QBF QBF AI safety via debate Geoffrey Irving Corresponding author: irving@openai.com Paul Christiano OpenAI Dario Amodei Abstract To make AI systems broadly useful for challenging real-world tasks, we need them to learn complex human goals and preferences. One approach to specifying complex goals asks humans to judge during training which agent behaviors are safe and useful, but this approach can fail if the task is too complicated for a human to directly judge. To help address this concern, we propose training agents via self play on a zero sum debate game.
Explore this link on the map →related reading
- [1805.00899] AI safety via debatearxiv.org
- The limits of AI safety via debate — LessWronglesswrong.com
- Optimal play in human-judged Debate usually won't answer your question — AI Alignment Forumalignmentforum.org
- An alignment safety case sketch based on debate — LessWronglesswrong.com
- Anthropic Fall 2023 Debate Progress Update — AI Alignment Forumalignmentforum.org
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- AI Debaters are More Persuasive when Arguing in Alignment with Their Own Beliefsarxiv.org
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Avoiding Obfuscation with Prover-Estimator Debatearxiv.org
- An AI alignment research agenda based on asymmetric debate and monitoring. — LessWronglesswrong.com
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- Where I agree and disagree with Eliezer — LessWronglesswrong.com