flâneur — a map of the web's best reading

Believe It or Not: How Deeply do LLMs Believe Implanted Facts?

alignment.anthropic.com · 1,479 words · saved by 3 readers

Techniques like synthetic document finetuning (SDF) have been proposed to modify the factual beliefs of LLMs. But do models genuinely believe these implanted facts? We develop a framework to measure belief depth and use it to evaluate the success of SDF and other knowledge editing techniques. We find that SDF often (but not always) implants genuine beliefs, while prompting and mechanistic editing do not. Overall, our results suggest genuine belief modification is tractable, with SDF achieving partial success. 📄 Paper, 💻 Code Research done as part of the Anthropic Fellows Program. The ability to control the factual beliefs of AI systems could be a useful tool for AI safety. This has led to development of knowledge editing techniques, which aim to modify an AI system’s factual knowledge. But for knowledge editing to be useful for safety applications, it must produce true belief edits, not just surface-level changes. In our paper, we develop a framework to measure belief depth: the degr

Believe It or Not: How Deeply do LLMs Believe Implanted Facts? Alignment Science Blog Believe It or Not: How Deeply do LLMs Believe Implanted Facts? Stewart Slocum 1 , October 21, 2025 Julian Minder 2,4 , Clement Dumas 3,4 , Henry Sleight 5 , Ryan Greenblatt 6 , Samuel Marks 7,† , Rowan Wang 7,† 1 Anthropic Fellows Program; 2 EPFL; 3 ENS Paris-Saclay, Université Paris-Saclay; 4 MATS; 5 Constellation; 6 Redwood Research; 7 Anthropic; † Equal contribution, order randomized tl;dr Techniques like synthetic document finetuning (SDF) have been proposed to modify the factual beliefs of LLMs. But do m

Explore this link on the map →

saved by

related reading