flâneur — a map of the web's best reading

ML Systems Will Have Weird Failure Modes

bounded-regret.ghost.io · 1,949 words · saved by 1 readers

Previously, I've argued that future ML systems might exhibit unfamiliar, emergent capabilities [https://bounded-regret.ghost.io/p/1527e9dd-c48d-4941-9b14-4f7293318d5c/], and that thought experiments provide one approach [https://bounded-regret.ghost.io/p/a2d733a7-108a-4587-97fb-db90f66ce030/] towards predicting these capabilities and their consequences. In this post I’ll describe a particular thought experiment in detail.

Previously, I've argued that future ML systems might exhibit unfamiliar, emergent capabilities , and that thought experiments provide one approach towards predicting these capabilities and their consequences. In this post I’ll describe a particular thought experiment in detail. We’ll see that taking thought experiments seriously often surfaces future risks that seem "weird" and alien from the point of view of current systems. I’ll also describe how I tend to engage with these thought experiments: I usually start out intuitively skeptical, but when I reflect on emergent behavior I find that som

Explore this link on the map →

related reading