AI Debaters are More Persuasive when Arguing in Alignment with Their Own Beliefs
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions.
\workshoptitle Multi-Turn Interactions in Large Language Models AI Debaters are More Persuasive when Arguing in Alignment with Their Own Beliefs María Victoria Carro 1,2 , Denise Alejandra Mester 1 1 1 footnotemark: 1 , Facundo Nieto 1,3 , Oscar Agustín Stanchi 5 , Guido Ernesto Bergman 4 , Mario Alejandro Leiva 6 , Eitan Sprejer 4 , Luca Nicolás Forziati Gangi 1 , Francisca Gauna Selasco 1 , Juan Gustavo Corvalán 1 , Gerardo I. Simari 6 , María Vanina Martinez 7 2 2 footnotemark: 2 1 FAIR, IALAB, Universidad de Buenos Aires, AR 2 Università degli Studi di Genova, IT 3 Universidad Nacional de
Explore this link on the map →related reading
- Anthropic Fall 2023 Debate Progress Update — AI Alignment Forumalignmentforum.org
- [1805.00899] AI safety via debatear5iv.labs.arxiv.org
- The Waluigi Effect (mega-post) — LessWronglesswrong.com
- BIG-bench/bigbench/benchmark_tasks/convinceme at main · google/BIG-bench · GitHubgithub.com
- Optimal play in human-judged Debate usually won't answer your question — AI Alignment Forumalignmentforum.org
- Measuring the Persuasiveness of Language Models \ Anthropicanthropic.com
- An alignment safety case sketch based on debate — LessWronglesswrong.com
- [2403.14380] On the Conversational Persuasiveness of Large Language Models: A Randomized Controlled Trialarxiv.org
- Externalized reasoning oversight: a research direction for language model alignment — AI Alignment Forumalignmentforum.org
- 39 - Evan Hubinger on Model Organisms of Misalignment | AXRP - the AI X-risk Research Podcastaxrp.net
- The limits of AI safety via debate — LessWronglesswrong.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com