Formalizing «Boundaries» with Markov blankets — LessWrong
[The post is largely a conceptual distillation of Andrew Critch’s Part 3a: Defining boundaries as directed Markov blankets.] By the end of this section, I want you to understand the following diagram (Pearlian causal diagram): Also: I will assume a basic familiarity with Markov chains in this post. First, I want you to imagine a simple Markov chain that represents the fact that a human influences itself over time: Second, I want you to imagine a Markov chain that represents the fact that the environment[1] influences itself over time: Okay. Now, notice that in between the human and its environment there’s some kind of boundary. For example, their skin (a physical boundary) and their interpretation/cognition (an informational boundary). If this were not a human but instead a bacterium, then the boundary I mean would (mostly) be the bacterium’s cell membrane. Third, imagine a Markov chain that represents that boundary influencing itself over time: Okay, so we have these three Markov chai
x Formalizing «Boundaries» with Markov blankets — LessWrong Boundaries (membranes) for AI safety by Chipmonk Boundaries / Membranes [technical] Causality Free Energy Principle Outer Alignment World Modeling Techniques AI World Modeling Frontpage 23 Formalizing «Boundaries» with Markov blankets by Chris Lakin 19th Sep 2023 4 min read 20 23 How could «boundaries» be formally specified? Markov blankets seem to be one fitting abstraction. [The post is largely a conceptual distillation of Andrew Critch’s Part 3a: Defining boundaries as directed Markov blankets .] Explaining Markov blankets By the e
Explore this link on the map →related reading
- «Boundaries», Part 3a: Defining boundaries as directed Markov blankets — AI Alignment Forumalignmentforum.org
- Understanding Agency through Markov Blankets — LessWronglesswrong.com
- A basic systems architecture for AI agents that do autonomous research — LessWronglesswrong.com
- Arjun Virkarjunvirk.com
- Scaling Managed Agents: Decoupling the brain from the hands \ Anthropicanthropic.com
- Neel Nanda on Mechanistic Interpretability: Progress, Limits, and Paths to Safer AI — EA Forumforum.effectivealtruism.org
- [1902.09469] Embedded Agencyarxiv.org
- Causal Markov condition - Wikipediaen.wikipedia.org
- How we contain Claude across products \ Anthropicanthropic.com
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- FLI AI Safety Research Landscape - Extended - v0.43futureoflife.org
- Causal scrubbing: Appendix — LessWronglesswrong.com