Real-Time Detection of Hallucinated Entities in Long-Form Generation
Large language models are now routinely used in high-stakes applications where hallucinations can cause serious harm, such as medical consultations or legal advice. Existing hallucination detection methods, however, are impractical for real-world use, as they are either limited to short factual queries or require costly external verification. We present a cheap, scalable method for real-time identification of hallucinated tokens in long-form generations, and scale it effectively to 70B parameter models. Our approach targets entity-level hallucinations—e.g., fabricated names, dates, citations—rather than claim-level, thereby naturally mapping to token-level labels and enabling streaming detection. We develop an annotation methodology that leverages web search to annotate model responses with grounded labels indicating which tokens correspond to fabricated entities. This dataset enables us to train effective hallucination classifiers with simple and efficient methods such as linear prob
Real-Time Detection of Hallucinated Entities in Long-Form Generation --> Real-Time Detection of Hallucinated Entities in Long-Form Generation Oscar Obeso * 1 , Andy Arditi * , Javier Ferrando, Joshua Freeman 1 Cameron Holmes 2 , Neel Nanda 1 ETH Zürich 2 MATS * Co-first authors Large language models are now routinely used in high-stakes applications where hallucinations can cause serious harm, such as medical consultations or legal advice. Existing hallucination detection methods, however, are impractical for real-world use, as they are either limited to short factual queries or require costly
Explore this link on the map →related reading
- Features as Rewards: Using Interpretability to Reduce Hallucinationsgoodfire.ai
- (Im)possibility of Automated Hallucination Detection in Large Language Modelsarxiv.org
- Hallucination Mitigation using Agentic AI Natural Language-Based Frameworksarxiv.org
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- HALVA: Hallucination Attenuated Language and Vision Assistantresearch.google
- Unfamiliar Finetuning Examples Control How Language Models Hallucinatearxiv.org
- Hallucination (artificial intelligence) - Wikipediaen.wikipedia.org
- The Dark Forest and Generative AImaggieappleton.com
- [2512.21577] A Unified Definition of Hallucination: It's The World Model, Stupid!arxiv.org
- Modifying LLM Beliefs with Synthetic Document Finetuningalignment.anthropic.com
- How confessions can keep language models honest | OpenAIopenai.com
- LLMs Know More Than What They Say - by Ruby Paiarjunbansal.substack.com