Believe It or Not: How Deeply do LLMs Believe Implanted Facts?
Techniques like synthetic document finetuning (SDF) have been proposed to modify the factual beliefs of LLMs. But do models genuinely believe these implanted facts? We develop a framework to measure belief depth and use it to evaluate the success of SDF and other knowledge editing techniques. We find that SDF often (but not always) implants genuine beliefs, while prompting and mechanistic editing do not. Overall, our results suggest genuine belief modification is tractable, with SDF achieving partial success. 📄 Paper, 💻 Code Research done as part of the Anthropic Fellows Program. The ability to control the factual beliefs of AI systems could be a useful tool for AI safety. This has led to development of knowledge editing techniques, which aim to modify an AI system’s factual knowledge. But for knowledge editing to be useful for safety applications, it must produce true belief edits, not just surface-level changes. In our paper, we develop a framework to measure belief depth: the degr
Believe It or Not: How Deeply do LLMs Believe Implanted Facts? Alignment Science Blog Believe It or Not: How Deeply do LLMs Believe Implanted Facts? Stewart Slocum 1 , October 21, 2025 Julian Minder 2,4 , Clement Dumas 3,4 , Henry Sleight 5 , Ryan Greenblatt 6 , Samuel Marks 7,† , Rowan Wang 7,† 1 Anthropic Fellows Program; 2 EPFL; 3 ENS Paris-Saclay, Université Paris-Saclay; 4 MATS; 5 Constellation; 6 Redwood Research; 7 Anthropic; † Equal contribution, order randomized tl;dr Techniques like synthetic document finetuning (SDF) have been proposed to modify the factual beliefs of LLMs. But do m
Explore this link on the map →saved by
related reading
- Modifying LLM Beliefs with Synthetic Document Finetuningalignment.anthropic.com
- Unfamiliar Finetuning Examples Control How Language Models Hallucinatearxiv.org
- The Waluigi Effect (mega-post) — LessWronglesswrong.com
- Standards for Belief Representations in LLMsarxiv.org
- Alignment faking in large language modelsarxiv.org
- BIG-bench/bigbench/benchmark_tasks/convinceme at main · google/BIG-bench · GitHubgithub.com
- Do language models possess knowledge (soundness)? - HackMDhackmd.io
- [2310.06824] The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasetsarxiv.org
- Quantifying Truesight With SAEs · Gwern.netgwern.net
- [2305.14251] FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generationarxiv.org
- How well do truth probes generalise? — LessWronglesswrong.com
- Thought Branches: Interpreting LLM Reasoning Requires Resamplingarxiv.org