Modifying LLM Beliefs with Synthetic Document Finetuning
alignment.anthropic.com · 6,160 words · saved by 7 readers
In this post, we study whether we can modify an LLM’s beliefs and investigate whether doing so could decrease risk from advanced AI systems.
Modifying LLM Beliefs with Synthetic Document Finetuning Alignment Science Blog Modifying LLM Beliefs with Synthetic Document Finetuning Rowan Wang April 24, 2025 Avery Griffin ‡ , Johannes Treutlein Ethan Perez, Julian Michael § , Fabien Roger, Sam Marks Anthropic; ‡ MATS; § Scale AI In this post, we study whether we can modify an LLM’s beliefs and investigate whether doing so could decrease risk from advanced AI systems. We describe a pipeline for modifying LLM beliefs via synthetic document finetuning and introduce a suite of evaluations that suggest our pipeline succeeds in inserting all b
saved by
related reading
- Believe It or Not: How Deeply do LLMs Believe Implanted Facts?alignment.anthropic.com
- [2510.17941] Believe It or Not: How Deeply do LLMs Believe Implanted Facts?arxiv.org
- Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- Student Projects - CS 2881R AI Safetyboazbk.github.io
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org
- [2603.07267] How to Steal Reasoning Without Reasoning Tracesarxiv.org
- Unfamiliar Finetuning Examples Control How Language Models Hallucinatearxiv.org
- [2605.13829] Negation Neglect: When models fail to learn negations in trainingarxiv.org
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Alignment faking in large language modelsarxiv.org
- 2308.03958arxiv.org
- [2506.17209] Fine-Tuning Lowers Safety and Disrupts Evaluation Consistencyarxiv.org