flâneur

truthfulQA_lin_evans.pdf

owainevans.github.io · 7,620 words · saved by 1 readers

N/A

TruthfulQA: Measuring How Models Mimic Human Falsehoods Stephanie Lin Jacob Hilton Owain Evans University of Oxford OpenAI University of Oxford sylin07@gmail.com jhilton@openai.com owaine@gmail.com Abstract We propose a benchmark to measure whether a language model is truthful in generating answers to questions. The benchmark comprises 817 questions that span 38…

saved by

related reading