Avoiding Obfuscation with Prover-Estimator Debate
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions.
Avoiding Obfuscation with Prover-Estimator Debate Avoiding Obfuscation with Prover-Estimator Debate Jonah Brown-Cohen Google DeepMind jonahbc@google.com &Geoffrey Irving UK AI Security Insitute geoffrey.irving@dsit.gov.uk &Georgios Piliouras Google DeepMind gpil@google.com Abstract Training powerful AI systems to exhibit desired behaviors hinges on the ability to provide accurate human supervision on increasingly complex tasks. A promising approach to this problem is to amplify human judgement by leveraging the power of two competing AIs in a debate about the correct solution to a given proble
Explore this link on the map →related reading
- Asymmetry of verification and verifier’s rule - Jason Weijasonwei.net
- [1805.00899] AI safety via debatear5iv.labs.arxiv.org
- An alignment safety case sketch based on debate — LessWronglesswrong.com
- The limits of AI safety via debate — LessWronglesswrong.com
- Anthropic Fall 2023 Debate Progress Update — AI Alignment Forumalignmentforum.org
- Zero Knowledge Proofs: An illustrated primer – A Few Thoughts on Cryptographic Engineeringblog.cryptographyengineering.com
- 2510.01346arxiv.org
- Optimal play in human-judged Debate usually won't answer your question — AI Alignment Forumalignmentforum.org
- AI Debaters are More Persuasive when Arguing in Alignment with Their Own Beliefsarxiv.org
- Robust Cooperation in the Prisoner's Dilemma — LessWronglesswrong.com
- ProofsArgsAndZK.pdfpeople.cs.georgetown.edu
- Computational Complexityblog.computationalcomplexity.org