flâneur — a map of the web's best reading

Unfamiliar Finetuning Examples Control How Language Models Hallucinate

arxiv.org · 8,938 words · saved by 1 readers

This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions.

Unfamiliar Finetuning Examples Control How Language Models Hallucinate Katie Kang 1 , Eric Wallace 1 , Claire Tomlin 1 , Aviral Kumar 2 , Sergey Levine 1 ( 1 UC Berkeley 2 Google DeepMind) Abstract Large language models are known to hallucinate when faced with unfamiliar queries, but the underlying mechanism that govern how models hallucinate are not yet fully understood. In this work, we find that unfamiliar examples in the models’ finetuning data – those that introduce concepts beyond the base model’s scope of knowledge – are crucial in shaping these errors. In particular, we find that an LL

Explore this link on the map →

saved by

related reading