[2506.13609] Avoiding Obfuscation with Prover-Estimator Debate
Abstract:Training powerful AI systems to exhibit desired behaviors hinges on the ability to provide accurate human supervision on increasingly complex tasks. A promising approach to this problem is to amplify human judgement by leveraging the power of two competing AIs in a debate about the correct solution to a given problem. Prior theoretical work has provided a complexity-theoretic formalization of AI debate, and posed the problem of designing protocols for AI debate that guarantee the correctness of human judgements for as complex a class of problems as possible. Recursive debates, in which debaters decompose a complex problem into simpler subproblems, hold promise for growing the class of problems that can be accurately judged in a debate. However, existing protocols for recursive debate run into the obfuscated arguments problem: a dishonest debater can use a computationally efficient strategy that forces an honest opponent to solve a computationally intractable problem to win. We mitigate this problem with a new recursive debate protocol that, under certain stability assumptions, ensures that an honest debater can win with a strategy requiring computational efficiency comparable to their opponent.
Avoiding Obfuscation with Prover-Estimator Debate Jonah Brown-Cohen Geoffrey Irving† Georgios Piliouras Google DeepMind Resolution Google DeepMind jonahbc@google.com irving@resolution.org gpil@google.com Lijie Chen Jiawei Li…
saved by
related reading
- Gradual Disempowerment from AI in Competitive Debatingdavidafrica.substack.com
- Avoiding Obfuscation with Prover-Estimator Debatearxiv.org
- [1805.00899] AI safety via debatear5iv.labs.arxiv.org
- confessions_paper.pdfcdn.openai.com
- [1805.00899] AI safety via debatearxiv.org
- The limits of AI safety via debate — LessWronglesswrong.com
- An alignment safety case sketch based on debate — LessWronglesswrong.com
- Asymmetry of verification and verifier’s rule - Jason Weijasonwei.net
- Anthropic Fall 2023 Debate Progress Update — AI Alignment Forumalignmentforum.org
- Deep Deceptiveness — LessWronglesswrong.com
- As Rocks May Think | Eric Jangevjang.com
- Optimal play in human-judged Debate usually won't answer your question — AI Alignment Forumalignmentforum.org