flâneur — a map of the web's best reading

World-Model Interpretability Is All We Need — AI Alignment Forum

alignmentforum.org · 8,560 words · saved by 1 readers

Imagine that we develop interpretability tools that allow us to flexibly understand and manipulate an AGI's world-model — but only its world-model. We would be able to see what the AGI knows, add or remove concepts from its mental ontology, and perhaps even use its world-model to run simulations/counterfactuals. But its thoughts and plans, and its hard-coded values and shards, would remain opaque to us. Would that be sufficient for robust alignment? I argue it would be. Primarily, this would solve the Pointers Problem. A central difficulty of alignment is that our values are functions of highly abstract variables, and that makes it hard to point an AI at them, instead of at easy-to-measure, shallow functions over sense-data. Cracking open a world-model would allow us to design metrics that have depth. From there, we'd have several ways to proceed: That leaves open the question of the "target metric". It primarily depends on what will be easy to specify — what concepts we'll find in the

x World-Model Interpretability Is All We Need — AI Alignment Forum AI Risk Interpretability (ML & AI) Natural Abstraction Research Agendas World Modeling Techniques AI Frontpage 11 World-Model Interpretability Is All We Need by Thane Ruthenis 14th Jan 2023 25 min read 22 11 Summary, by sections : Perfect world-model interpretability seems both sufficient for robust alignment (via a decent variety of approaches) and realistically attainable (compared to "perfect interpretability" in general, i. e. insight into AIs' heuristics, goals, and thoughts as well). Main arguments: the NAH + internal int

Explore this link on the map →

related reading