flâneur — a map of the web's best reading

Modifying LLM Beliefs with Synthetic Document Finetuning

alignment.anthropic.com · 6,160 words · saved by 6 readers

In this post, we study whether we can modify an LLM’s beliefs and investigate whether doing so could decrease risk from advanced AI systems.

Modifying LLM Beliefs with Synthetic Document Finetuning Alignment Science Blog Modifying LLM Beliefs with Synthetic Document Finetuning Rowan Wang April 24, 2025 Avery Griffin ‡ , Johannes Treutlein Ethan Perez, Julian Michael § , Fabien Roger, Sam Marks Anthropic; ‡ MATS; § Scale AI In this post, we study whether we can modify an LLM’s beliefs and investigate whether doing so could decrease risk from advanced AI systems. We describe a pipeline for modifying LLM beliefs via synthetic document finetuning and introduce a suite of evaluations that suggest our pipeline succeeds in inserting all b

Explore this link on the map →

saved by

related reading