[2510.17941] Believe It or Not: How Deeply do LLMs Believe Implanted Facts?
Abstract:Knowledge editing techniques promise to implant new factual knowledge into large language models (LLMs). But do LLMs really believe these facts? We develop a framework to measure belief depth and use it to evaluate the success of knowledge editing techniques. We operationalize belief depth as the extent to which implanted knowledge 1) generalizes to related contexts (e.g. Fermi estimates several logical steps removed), 2) is robust to self-scrutiny and direct challenge, and 3) is represented similarly to genuine knowledge (as measured by linear probes). Our evaluations show that simple prompting and mechanistic editing techniques fail to implant knowledge deeply. In contrast, Synthetic Document Finetuning (SDF) - where models are trained on LLM-generated documents consistent with a fact - often succeeds at implanting beliefs that behave similarly to genuine knowledge. However, SDF's success is not universal, as implanted beliefs that contradict basic world knowledge are brittle and representationally distinct from genuine knowledge. Overall, our work introduces measurable criteria for belief depth and enables the rigorous evaluation necessary for deploying knowledge editing in real-world applications.
B ELIEVE I T OR N OT: H OW D EEPLY DO LLM S B ELIEVE I MPLANTED FACTS ? Stewart Slocum1 Julian Minder2,4 , Clement Dumas3,4 Henry Sleight5 , Ryan Greenblatt6 , Samuel Marks7,† , Rowan Wang7,† 1 Anthropic Fellows Program, 2 EPFL, 3 ENS Paris-Saclay, Université Paris-Saclay, 4 MATS, 5…
saved by
related reading
- Believe It or Not: How Deeply do LLMs Believe Implanted Facts?alignment.anthropic.com
- [2503.03750] The MASK Benchmark: Disentangling Honesty From Accuracy in AI Systemsarxiv.org
- Modifying LLM Beliefs with Synthetic Document Finetuningalignment.anthropic.com
- [2605.27288] It's Not Always Sycophancy: Measuring LLM Conformity as a Function of Epistemic Uncertaintyarxiv.org
- [2605.13829] Negation Neglect: When models fail to learn negations in trainingarxiv.org
- Unfamiliar Finetuning Examples Control How Language Models Hallucinatearxiv.org
- Standards for Belief Representations in LLMsarxiv.org
- Large Language Model: world models or surface statistics?thegradient.pub
- Do language models possess knowledge (soundness)? - HackMDhackmd.io
- Self-Adapting Language Modelsarxiv.org
- 2202.05262arxiv.org
- Can a Language Model Learn Facts Continually in Its Weights?labs.baseten.co