World-Model Interpretability Is All We Need — AI Alignment Forum
Imagine that we develop interpretability tools that allow us to flexibly understand and manipulate an AGI's world-model — but only its world-model. We would be able to see what the AGI knows, add or remove concepts from its mental ontology, and perhaps even use its world-model to run simulations/counterfactuals. But its thoughts and plans, and its hard-coded values and shards, would remain opaque to us. Would that be sufficient for robust alignment? I argue it would be. Primarily, this would solve the Pointers Problem. A central difficulty of alignment is that our values are functions of highly abstract variables, and that makes it hard to point an AI at them, instead of at easy-to-measure, shallow functions over sense-data. Cracking open a world-model would allow us to design metrics that have depth. From there, we'd have several ways to proceed: That leaves open the question of the "target metric". It primarily depends on what will be easy to specify — what concepts we'll find in the
x World-Model Interpretability Is All We Need — AI Alignment Forum AI Risk Interpretability (ML & AI) Natural Abstraction Research Agendas World Modeling Techniques AI Frontpage 11 World-Model Interpretability Is All We Need by Thane Ruthenis 14th Jan 2023 25 min read 22 11 Summary, by sections : Perfect world-model interpretability seems both sufficient for robust alignment (via a decent variety of approaches) and realistically attainable (compared to "perfect interpretability" in general, i. e. insight into AIs' heuristics, goals, and thoughts as well). Main arguments: the NAH + internal int
Explore this link on the map →related reading
- World Models: Computing the Uncomputablenotboring.co
- Dario Amodei — The Urgency of Interpretabilitydarioamodei.com
- Language Models, World Models, and Human Model-Buildinglingo.csail.mit.edu
- LLMs and World Models, Part 1 - by Melanie Mitchellaiguide.substack.com
- Research Agenda: Synthesizing Standalone World-Models — AI Alignment Forumalignmentforum.org
- How Can Interpretability Researchers Help AGI Go Well? — AI Alignment Forumalignmentforum.org
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- A Functional Taxonomy of World Models - Dr. Fei-Fei Lidrfeifei.substack.com
- pdfopenreview.net
- Neel Nanda on Mechanistic Interpretability: Progress, Limits, and Paths to Safer AI — EA Forumforum.effectivealtruism.org
- On Optimism for Interpretabilitygoodfire.ai
- Building AGI Using Language Models | Leo Gaobmk.sh