The two types of LLM preferences
The main issue with every single experiment of this sort is that the results are not robust to reasonable variations in the prompt. The LLM’s decisions usually vary a lot based on factors that we do not consider meaningful; in other words, they are inconsistent. I’ve observed prompt-driven preference variability many times myself, but the paper people cite for this nowadays is Randomness, Not Representation: The Unreliability of Evaluating Cultural Alignment in LLMs (Khan, Casper, Hadfield-Menell, 2025). I feel there is an ontological issue deep at play. We don’t actually know what we are talking about when we measure LLM values and preferences; or how far these words are from their meaning when applied to people. In particular, I want to highlight that there is a spectrum of preferences between: strong preferences: preferences that persist across reasonable variations in context, wording, and framing; weak preferences: statistical tendencies that show up when averaged across many tria
The standard approach to measure values or preferences1 of LLMs is to: construct binary questions that would reflect a preference when posed to a person; pose many such questions to an LLM; statistically analyze the responses to find legible preferences. The main issue with every single experiment of this sort is that the results are not robust to reasonable variations in the prompt. The LLM’s decisions usually vary a lot based on factors that we do not consider meaningful; in other words, they are inconsistent. I’ve observed prompt-driven preference variability many times myself, but…
related reading
- Probing Persona-Dependent Preferences in Language Modelsarxiv.org
- Guardian Angels: LLM Personalization for Productivity and Security · Gwern.netgwern.net
- [2506.00751] Alignment Revisited: Are Large Language Models Consistent in Stated and Revealed Preferences?arxiv.org
- Text role playing games to discover task preferences in LLMssimonlermen.substack.com
- What We’ve Learned From A Year of Building with LLMs – Applied LLMsapplied-llms.org
- [2503.10990] Statistical Impossibility and Possibility of Aligning LLMs with Human Preferences: From Condorcet Paradox to Nash Equilibriumarxiv.org
- Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts | RLHFlowrlhflow.github.io
- 2410.12851arxiv.org
- [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Modelarxiv.org
- [2604.21751] Why are all LLMs Obsessed with Japanese Culture? On the Hidden Cultural and Regional Biases of LLMsarxiv.org
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- 2307.12950.pdfarxiv.org