Unfamiliar Finetuning Examples Control How Language Models Hallucinate
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions.
Unfamiliar Finetuning Examples Control How Language Models Hallucinate Katie Kang 1 , Eric Wallace 1 , Claire Tomlin 1 , Aviral Kumar 2 , Sergey Levine 1 ( 1 UC Berkeley 2 Google DeepMind) Abstract Large language models are known to hallucinate when faced with unfamiliar queries, but the underlying mechanism that govern how models hallucinate are not yet fully understood. In this work, we find that unfamiliar examples in the models’ finetuning data – those that introduce concepts beyond the base model’s scope of knowledge – are crucial in shaping these errors. In particular, we find that an LL
Explore this link on the map →saved by
related reading
- Believe It or Not: How Deeply do LLMs Believe Implanted Facts?alignment.anthropic.com
- Modifying LLM Beliefs with Synthetic Document Finetuningalignment.anthropic.com
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- Features as Rewards: Using Interpretability to Reduce Hallucinationsgoodfire.ai
- Patterns for Building LLM-based Systems & Productseugeneyan.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org
- Fine-Tuning Llama-2: Tailoring Models to Unique Applicationsanyscale.com
- [2203.02155] Training language models to follow instructions with human feedbackarxiv.org
- Alignment faking in large language modelsarxiv.org
- The State of Reinforcement Learning for LLM Reasoningmagazine.sebastianraschka.com
- rl-for-llms.md · GitHubgist.github.com