Goodhart's Law - AI Alignment Forum
Goodhart's Law states that when a proxy for some value becomes the target of optimization pressure, the proxy will cease to be a good proxy. One form of Goodhart is demonstrated by the Soviet story of a factory graded on how many shoes they produced (a good proxy for productivity) – they soon began producing a higher number of tiny shoes. Useless, but the numbers look good. Goodhart's Law is of particular relevance to AI Alignment. Suppose you have something which is generally a good proxy for "the stuff that humans care about", it would be dangerous to have a powerful AI optimize for the proxy, in accordance with Goodhart's law, the proxy will breakdown. GOODHART TAXONOMY In Goodhart Taxonomy, Scott Garrabrant identifies four kinds of Goodharting: * Regressional Goodhart - When selecting for a proxy measure, you select not only for the true goal, but also for the difference between the proxy and the goal. * Causal Goodhart - When there is a non-causal correlation between the proxy and the goal, intervening on the proxy may fail to intervene on the goal. * Extremal Goodhart - Worlds in which the proxy takes an extreme value may be very different from the ordinary worlds in which the correlation between the proxy and the goal was observed. * Adversarial Goodhart - When you optimize for a proxy, you provide an incentive for adversaries to correlate their goal with your proxy, thus destroying the correlation with your goal. SEE ALSO * Groupthink, Information cascade, Affective death spiral * Adaptation executers, Superstimulus * Signaling, Filtered evidence * Cached thought * Modesty argument, Egalitarianism * Rationalization, Dark arts * Epistemic hygiene * Scoring rule
x Goodhart's Law — AI Alignment Forum Goodhart's Law Edited by Ruby , Vladimir_Nesov , Xylix , et al. last updated 14th Jul 2026 Goodhart's Law states that when a proxy for some value becomes the target of optimization pressure, the proxy will cease to be a good proxy. One form of Goodhart is demonstrated by the Soviet story of a factory graded on how many shoes they produced (a good proxy for productivity) – they soon began producing a higher number of tiny shoes. Useless, but the numbers look good. Goodhart's Law is of particular relevance to AI Alignment . Suppose you have something which i
Explore this link on the map →related reading
- Goodhart's law - Wikipediaen.wikipedia.org
- Too much efficiency makes everything worse: overfitting and the strong version of Goodhart’s law | Jascha’s blogsohl-dickstein.github.io
- Goodhart's Law Isn't as Useful as You Might Think - Commoncogcommoncog.com
- Why Agent Foundations? An Overly Abstract Explanation — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Goodhart's Law and Why Measurement is Hard — Ribbonfarmribbonfarm.com
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- Bayesianism versus conservatism versus Goodhart — AI Alignment Forumalignmentforum.org
- Why the tails (sometimes) don’t come apartuniversalprior.substack.com