Utility Maximization = Description Length Minimization — LessWrong
There’s a useful intuitive notion of “optimization” as pushing the world into a small set of states, starting from any of a large number of states. Visually: Yudkowsky and Flint both have notable formalizations of this “optimization as compression” idea. This post presents a formalization of optimization-as-compression grounded in information theory. Specifically: to “optimize” a system is to reduce the number of bits required to represent the system state using a particular encoding. In other words, “optimizing” a system means making it compressible (in the information-theoretic sense) by a particular model. This formalization turns out to be equivalent to expected utility maximization, and allows us to interpret any expected utility maximizer as “trying to make the world look like a particular model”. Before diving into the formalism, we’ll walk through a conceptual example, taken directly from Flint’s Ground of Optimization: building a house. Here’s Flint’s diagram: The key idea her
x Utility Maximization = Description Length Minimization — LessWrong Basic Foundations for Agent Models Information theory Optimization Utility Functions AI Rationality Curated 225 Utility Maximization = Description Length Minimization by johnswentworth 18th Feb 2021 AI Alignment Forum 7 min read 54 225 Ω 72 There’s a useful intuitive notion of “optimization” as pushing the world into a small set of states, starting from any of a large number of states. Visually: Yudkowsky and Flint both have notable formalizations of this “optimization as compression” idea. This post presents a formalization
related reading
- Visual Information Theory -- colah's blogcolah.github.io
- Compression and Intelligencegreene.sh
- Shtetl-Optimized >> Blog Archive >> The First Law of Complexodynamicsscottaaronson.blog
- Six (and a half) intuitions for KL divergence — LessWronglesswrong.com
- Optimality is the tiger, and agents are its teeth — LessWronglesswrong.com
- [2301.12987] The Optimal Choice of Hypothesis Is the Weakest, Not the Shortestarxiv.org
- A Mathematical Theory of Communicationpeople.math.harvard.edu
- Compression is predictionngrok.com
- Reward is not the optimization target — LessWronglesswrong.com
- Data Compression Explainedmattmahoney.net
- Visual Information Theory -- colah's blogcolah.github.io
- [2206.07867] A visual introduction to information theoryarxiv.org