Bayesianism versus conservatism versus Goodhart — AI Alignment Forum
Key argument: if we use a non-Bayesian conservative approach, such as a minimum over different utility functions, then we better have a good reason a…
x Bayesianism versus conservatism versus Goodhart — AI Alignment Forum Goodhart's Law AI Frontpage 8 Bayesianism versus conservatism versus Goodhart by Stuart_Armstrong 16th Jul 2021 7 min read 2 8 Key argument: if we use a non-Bayesian conservative approach, such as a minimum over different utility functions, then we better have a good reason as to why that would work. But if we have that reason, we can use it to make the whole thing into a Bayesian mix, which can also allow us to trade off that advantage against other possible gains. I've defended using Bayesian averaging of possible utility
Explore this link on the map →related reading
- Approximately Bayesian Reasoning: Knightian Uncertainty, Goodhart, and the Look-Elsewhere Effect — LessWronglesswrong.com
- Goodhart's law - Wikipediaen.wikipedia.org
- Too much efficiency makes everything worse: overfitting and the strong version of Goodhart’s law | Jascha’s blogsohl-dickstein.github.io
- Bayesian Mindsetcold-takes.com
- On The Independence Axiom — LessWronglesswrong.com
- Why The Focus on Expected Utility Maximisers? — LessWronglesswrong.com
- Goodhart's Law — AI Alignment Forumalignmentforum.org
- Probabilities Without Measurements | Vaden Masranivmasrani.github.io
- The Bayes Banditfrancesco215.github.io
- Efficient tradeoffs and the safety-usefulness tradeoff model — LessWronglesswrong.com
- Why We Can't Take Expected Value Estimates Literally (Even When They're Unbiased) — LessWronglesswrong.com
- Why we can't take expected value estimates literally (even when they're unbiased) - The GiveWell Blogblog.givewell.org