AI Debaters are More Persuasive when Arguing in Alignment with Their Own Beliefs
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions.
\workshoptitle Multi-Turn Interactions in Large Language Models AI Debaters are More Persuasive when Arguing in Alignment with Their Own Beliefs María Victoria Carro 1,2 , Denise Alejandra Mester 1 1 1 footnotemark: 1 , Facundo Nieto 1,3 , Oscar Agustín Stanchi 5 , Guido Ernesto Bergman 4 , Mario Alejandro Leiva 6 , Eitan Sprejer 4 , Luca Nicolás Forziati Gangi 1 , Francisca Gauna Selasco 1 , Juan Gustavo Corvalán 1 , Gerardo I. Simari 6 , María Vanina Martinez 7 2 2 footnotemark: 2 1 FAIR, IALAB, Universidad de Buenos Aires, AR 2 Università degli Studi di Genova, IT 3 Universidad Nacional de
related reading
- Anthropic Fall 2023 Debate Progress Update — AI Alignment Forumalignmentforum.org
- [1805.00899] AI safety via debatear5iv.labs.arxiv.org
- Alignment is not solved but it increasingly looks solvablealigned.substack.com
- BIG-bench/bigbench/benchmark_tasks/convinceme at main · google/BIG-bench · GitHubgithub.com
- The limits of AI safety via debate — LessWronglesswrong.com
- [1805.00899] AI safety via debatearxiv.org
- [2506.13609] Avoiding Obfuscation with Prover-Estimator Debatearxiv.org
- LMCA_dataset.pdfandrew.cmu.edu
- Optimal play in human-judged Debate usually won't answer your question — AI Alignment Forumalignmentforum.org
- Deep Deceptiveness — LessWronglesswrong.com
- Measuring the Persuasiveness of Language Models \ Anthropicanthropic.com
- Gradual Disempowerment from AI in Competitive Debatingdavidafrica.substack.com