2023.emnlp-main.330.pdf
aclanthology.org · 5,223 words · saved by 1 readers
N/A
Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback Katherine Tian,∗† Eric Mitchell,∗‡ Allan Zhou,‡ Archit Sharma,‡ Rafael Rafailov‡ Huaxiu Yao,‡ Chelsea Finn,‡ Christopher D. Manning‡ † Harvard University ‡ Stanford University ktian@college.harvard.edu eric.mitchell@cs.stanford.edu…
related reading
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- Large Language Models Must Be Taught to Know What They Don't Knowarxiv.org
- gpt-4.pdfcdn.openai.com
- Cycles of Thought: Measuring LLM Confidence through Stable Explanationsarxiv.org
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- [1706.04599] On Calibration of Modern Neural Networksarxiv.org
- [2405.20974] SaySelf: Teaching LLMs to Express Confidence with Self-Reflective Rationalesarxiv.org
- [2003.07892] Calibration of Pre-trained Transformersarxiv.org
- Language Models Learn to Mislead Humans via RLHFarxiv.org
- Unfamiliar Finetuning Examples Control How Language Models Hallucinatearxiv.org
- [2203.02155] Training language models to follow instructions with human feedbackarxiv.org
- [2207.05221] Language Models (Mostly) Know What They Knowarxiv.org