flâneur — a map of the web's best reading

The two types of LLM preferences

newsletter.danielpaleka.com · saved by 1 readers

The main issue with every single experiment of this sort is that the results are not robust to reasonable variations in the prompt. The LLM’s decisions usually vary a lot based on factors that we do not consider meaningful; in other words, they are inconsistent. I’ve observed prompt-driven preference variability many times myself, but the paper people cite for this nowadays is Randomness, Not Representation: The Unreliability of Evaluating Cultural Alignment in LLMs (Khan, Casper, Hadfield-Menell, 2025). I feel there is an ontological issue deep at play. We don’t actually know what we are talking about when we measure LLM values and preferences; or how far these words are from their meaning when applied to people. In particular, I want to highlight that there is a spectrum of preferences between: strong preferences: preferences that persist across reasonable variations in context, wording, and framing; weak preferences: statistical tendencies that show up when averaged across many tria

The main issue with every single experiment of this sort is that the results are not robust to reasonable variations in the prompt. The LLM’s decisions usually vary a lot based on factors that we do not consider meaningful; in other words, they are inconsistent. I’ve observed prompt-driven preference variability many times myself, but the paper people cite for this nowadays is Randomness, Not Representation: The Unreliability of Evaluating Cultural Alignment in LLMs (Khan, Casper, Hadfield-Menell, 2025). I feel there is an ontological issue deep at play. We don’t actually know what we are talk

Explore this link on the map →