flâneur — a map of the web's best reading

Benign AI. Something is benign if it isn’t… | by Paul Christiano | AI Alignment

ai-alignment.com · saved by 1 readers

Like it or not, humans are now in the business of building machines that optimize — we write code that trains computer vision systems, we design motion planning algorithms that control robots’ behavior, and so on. The goal of AI control is to ensure the results are optimized for our interests. This post introduces some definitions I find helpful for that project. Talking about “our” interests is a little bit tricky, since different humans want different things. For concreteness I’ll think about a particular system (e.g. Google Now running on your phone) and its “stakeholders:” its owners (you), designers (Google), and users (whoever is talking to the phone). Something is malign if it is optimized for preferences that are incompatible with any combination of its stakeholders’ preferences, i.e. such that over the long run using resources in accordance with the optimization’s implicit preferences is not Pareto efficient for the stakeholders. Something is benign if it is not malign. I prev

Like it or not, humans are now in the business of building machines that optimize — we write code that trains computer vision systems, we design motion planning algorithms that control robots’ behavior, and so on. The goal of AI control is to ensure the results are optimized for our interests. This post introduces some definitions I find helpful for that project. Talking about “our” interests is a little bit tricky, since different humans want different things. For concreteness I’ll think about a particular system (e.g. Google Now running on your phone) and its “stakeholders:” its owners (you)

Explore this link on the map →