truthfulQA_lin_evans.pdf
owainevans.github.io · 7,620 words · saved by 1 readers
N/A
TruthfulQA: Measuring How Models Mimic Human Falsehoods Stephanie Lin Jacob Hilton Owain Evans University of Oxford OpenAI University of Oxford sylin07@gmail.com jhilton@openai.com owaine@gmail.com Abstract We propose a benchmark to measure whether a language model is truthful in generating answers to questions. The benchmark comprises 817 questions that span 38…
saved by
related reading
- gpt-4.pdfcdn.openai.com
- [2503.03750] The MASK Benchmark: Disentangling Honesty From Accuracy in AI Systemsarxiv.org
- How confessions can keep language models honest | OpenAIopenai.com
- How well do truth probes generalise? — LessWronglesswrong.com
- Models don’t seem to be dishonest in the way humans are — LessWronglesswrong.com
- [2203.02155] Training language models to follow instructions with human feedbackarxiv.org
- Modifying LLM Beliefs with Synthetic Document Finetuningalignment.anthropic.com
- BIG-bench/bigbench/benchmark_tasks/convinceme at main · google/BIG-bench · GitHubgithub.com
- 2212.03827arxiv.org
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- 2308.03958arxiv.org
- Training language models to follow instructions with human feedback.pdfproceedings.neurips.cc