flâneur — a map of the web's best reading

Avoiding Obfuscation with Prover-Estimator Debate

arxiv.org · 33,629 words · saved by 1 readers

This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions.

Avoiding Obfuscation with Prover-Estimator Debate Avoiding Obfuscation with Prover-Estimator Debate Jonah Brown-Cohen Google DeepMind jonahbc@google.com &Geoffrey Irving UK AI Security Insitute geoffrey.irving@dsit.gov.uk &Georgios Piliouras Google DeepMind gpil@google.com Abstract Training powerful AI systems to exhibit desired behaviors hinges on the ability to provide accurate human supervision on increasingly complex tasks. A promising approach to this problem is to amplify human judgement by leveraging the power of two competing AIs in a debate about the correct solution to a given proble

Explore this link on the map →

related reading