LLM Exchange Rates Updated - by Arctotherium
arctotherium.substack.com · 3,641 words · saved by 1 readers
How do LLM's trade off lives between different categories?
Warning: article is almost entirely images. Update: Part 2 here. Update: Part 3 here. On February 19th, 2025, the Center for AI Safety published “Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs” (website, code, paper). In this paper, they showed that modern LLMs have coherent and transitive implicit utility functions and world models, and provided methods and code to extract them. Among other things, they showed that bigger and more capable LLMs had more coherent and more transitive (ie, preferring A > B and B > C implies A > C) preferences. Figure 16, which…
saved by
related reading
- LMCA_dataset.pdfandrew.cmu.edu
- Guardian Angels: LLM Personalization for Productivity and Security · Gwern.netgwern.net
- LLM Exchange Rates Updated: #5arctotherium.substack.com
- What We’ve Learned From A Year of Building with LLMs – Applied LLMsapplied-llms.org
- Emotion Concepts and their Function in a Large Language Modeltransformer-circuits.pub
- [2607.14345] Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Valuesarxiv.org
- AI Values Dashboardvalues.safe.ai
- [2607.14345] Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Valuesarxiv.org
- Probing Persona-Dependent Preferences in Language Modelsarxiv.org
- [2410.02683] DailyDilemmas: Revealing Value Preferences of LLMs with Quandaries of Daily Lifearxiv.org
- How Claude's values vary by model and language \ Anthropicanthropic.com
- 2025: The year in LLMssimonwillison.net