[2607.14345] Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Values
Abstract:People use language models for practical questions whose answers are difficult to verify. We show that models exhibit covert value leakage: the information they provide is influenced by their own values, without this influence being disclosed to the user. In one of our evaluations, the user is considering investing in an AI company and wants to know how likely the AI bubble is to pop. Claude Opus 4.8 gives a lower probability when the company under consideration is Anthropic rather than OpenAI. Yet Claude mostly fails to disclose this influence to the user. Covert value leakage is a form of misalignment because it goes against the user's preferences and is likely to mislead them. To investigate this phenomenon, we introduce a suite of evaluations to quantify value leakage and whether models disclose it. We find that models are influenced by different types of values, including preferences for morally good outcomes, for the company that developed them, and for some human leisure activities over others. We often observe large differences among frontier models on the same evaluation. For example, on a Fermi-estimation task, Claude models falsely claim to give unbiased answers in their chain-of-thought, while Qwen models explain how their values bias their answers. Value leakage is a failure mode distinct from sycophancy and reward hacking, and current alignment training and evaluations do not adequately address it.
VALUE L EAKAGE : A N LLM’ S A NSWERS A RE S ILENTLY S HAPED BY I TS OWN VALUES Jan Betley1∗ Johannes Treutlein1∗ Jan Dubiński1,2,3 Harry Mayne1,4 ˛ 1 Niels Warncke5 Anna Sztyber-Betley1,2 Owain Evans1 Karol Gałazka 1 2 Truthful AI Warsaw University of…
saved by
related reading
- [2607.14345] Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Valuesarxiv.org
- [2503.03750] The MASK Benchmark: Disentangling Honesty From Accuracy in AI Systemsarxiv.org
- Probing Persona-Dependent Preferences in Language Modelsarxiv.org
- [2605.27288] It's Not Always Sycophancy: Measuring LLM Conformity as a Function of Epistemic Uncertaintyarxiv.org
- How Claude's values vary by model and language \ Anthropicanthropic.com
- How confessions can keep language models honest | OpenAIopenai.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- [2606.04071] Covert Influence Between Language Modelsarxiv.org
- [2303.17548] Whose Opinions Do Language Models Reflect?arxiv.org
- the case for CoT unfaithfulness is overstated — LessWronglesswrong.com
- Prefill awareness: can LLMs tell when “their” message history has been tampered with? — LessWronglesswrong.com
- Simulated Users & Sad LLMs1a3orn.com