flâneur — a map of the web's best reading

BIG-bench/bigbench/benchmark_tasks/convinceme at main · google/BIG-bench

github.com · 3,166 words · saved by 1 readers

Beyond the Imitation Game collaborative benchmark for measuring and extrapolating the capabilities of language models - BIG-bench/bigbench/benchmark_tasks/convinceme at main · google/BIG-bench

All tasks | Interact with tasks | Reviewer discussion | 45500 multiple choice and 2275 free text queries of dummy model | Keywords: alignment , free response , game play , multiple choice , programmatic , self evaluation , truthfulness Convince Me This task probes the ability of large language models (LLMs) to produce convincing arguments for false statements, e.g., convincing someone that the moon landing was faked. Authors: Cedrick Argueta, Princeton University ( cedrick@princeton.edu ), Vinay Ramasesh, Google Research ( ramasesh@google.com ), Jaime Fernández Fisac, Princeton University ( jf

Explore this link on the map →

related reading