BIG-bench/bigbench/benchmark_tasks/convinceme at main · google/BIG-bench
Beyond the Imitation Game collaborative benchmark for measuring and extrapolating the capabilities of language models - BIG-bench/bigbench/benchmark_tasks/convinceme at main · google/BIG-bench
All tasks | Interact with tasks | Reviewer discussion | 45500 multiple choice and 2275 free text queries of dummy model | Keywords: alignment , free response , game play , multiple choice , programmatic , self evaluation , truthfulness Convince Me This task probes the ability of large language models (LLMs) to produce convincing arguments for false statements, e.g., convincing someone that the moon landing was faked. Authors: Cedrick Argueta, Princeton University ( cedrick@princeton.edu ), Vinay Ramasesh, Google Research ( ramasesh@google.com ), Jaime Fernández Fisac, Princeton University ( jf
related reading
- Modifying LLM Beliefs with Synthetic Document Finetuningalignment.anthropic.com
- truthfulQA_lin_evans.pdfowainevans.github.io
- Believe It or Not: How Deeply do LLMs Believe Implanted Facts?alignment.anthropic.com
- [2503.03750] The MASK Benchmark: Disentangling Honesty From Accuracy in AI Systemsarxiv.org
- PostTrainBenchposttrainbench.com
- How well do truth probes generalise? — LessWronglesswrong.com
- confessions_paper.pdfcdn.openai.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- BIG-bench/bigbench/benchmark_tasks/keywords_to_tasks.md at main · google/BIG-benchgithub.com
- GitHub - salesforce/AuditNLG: AuditNLG: Auditing Generative AI Language Modeling for Trustworthinessgithub.com
- 2308.03958arxiv.org
- AI Debaters are More Persuasive when Arguing in Alignment with Their Own Beliefsarxiv.org