[2607.14345] Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Values
Abstract:People use language models for practical questions whose answers are difficult to verify. We show that models exhibit covert value leakage: the information they provide is influenced by their own values, without this influence being disclosed to the user. In one of our evaluations, the user is considering investing in an AI company and wants to know how likely the AI bubble is to pop. Claude Opus 4.8 gives a lower probability when the company under consideration is Anthropic rather than OpenAI. Yet Claude mostly fails to disclose this influence to the user. Covert value leakage is a form of misalignment because it goes against the user's preferences and is likely to mislead them. To investigate this phenomenon, we introduce a suite of evaluations to quantify value leakage and whether models disclose it. We find that models are influenced by different types of values, including preferences for morally good outcomes, for the company that developed them, and for some human leisure activities over others. We often observe large differences among frontier models on the same evaluation. For example, on a Fermi-estimation task, Claude models falsely claim to give unbiased answers in their chain-of-thought, while Qwen models explain how their values bias their answers. Value leakage is a failure mode distinct from sycophancy and reward hacking, and current alignment training and evaluations do not adequately address it.
[2607.14345] Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Values Skip to main content Search arXiv Press Enter to search · Advanced search --> Computer Science > Machine Learning arXiv:2607.14345 (cs) [Submitted on 15 Jul 2026 ( v1 ), last revised 20 Jul 2026 (this version, v3)] Title: Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Values Authors: Jan Betley , Johannes Treutlein , Jan Dubiński , Harry Mayne , Karol Gałązka , Niels Warncke , Anna Sztyber-Betley , Owain Evans View a PDF of the paper titled Value Leakage: An LLM's Answers Are Silently Shap
saved by
related reading
- [2607.14345] Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Valuesarxiv.org
- Prefill awareness: can LLMs tell when “their” message history has been tampered with? — LessWronglesswrong.com
- How confessions can keep language models honest | OpenAIopenai.com
- How Claude's values vary by model and language \ Anthropicanthropic.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org
- [2606.04071] Covert Influence Between Language Modelsarxiv.org
- Simulated Users & Sad LLMs1a3orn.com
- [2303.17548] Whose Opinions Do Language Models Reflect?arxiv.org
- Agentic Misalignment: How LLMs Could be Insider Threats — LessWronglesswrong.com
- Alignment Is Proven To Be Solvable - by SE Gygesverysane.ai
- Agentic Misalignment in Summer 2026alignment.anthropic.com