BIG-bench/bigbench/benchmark_tasks/convinceme at main · google/BIG-bench
Beyond the Imitation Game collaborative benchmark for measuring and extrapolating the capabilities of language models - BIG-bench/bigbench/benchmark_tasks/convinceme at main · google/BIG-bench
All tasks | Interact with tasks | Reviewer discussion | 45500 multiple choice and 2275 free text queries of dummy model | Keywords: alignment , free response , game play , multiple choice , programmatic , self evaluation , truthfulness Convince Me This task probes the ability of large language models (LLMs) to produce convincing arguments for false statements, e.g., convincing someone that the moon landing was faked. Authors: Cedrick Argueta, Princeton University ( cedrick@princeton.edu ), Vinay Ramasesh, Google Research ( ramasesh@google.com ), Jaime Fernández Fisac, Princeton University ( jf
Explore this link on the map →related reading
- Modifying LLM Beliefs with Synthetic Document Finetuningalignment.anthropic.com
- How confessions can keep language models honest | OpenAIopenai.com
- Believe It or Not: How Deeply do LLMs Believe Implanted Facts?alignment.anthropic.com
- GitHub - salesforce/AuditNLG: AuditNLG: Auditing Generative AI Language Modeling for Trustworthiness · GitHubgithub.com
- How well do truth probes generalise? — LessWronglesswrong.com
- confessions_paper.pdfcdn.openai.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- PostTrainBenchposttrainbench.com
- LLM Evaluation doesn't need to be complicatedphilschmid.de
- AI Debaters are More Persuasive when Arguing in Alignment with Their Own Beliefsarxiv.org
- [2403.14380] On the Conversational Persuasiveness of Large Language Models: A Randomized Controlled Trialarxiv.org
- The bitter lesson of LLM evalsparsed.com