Working through a small tiling result — LessWrong
tl;dr it seems that you can get basic tiling to work by proving that there will be safety proofs in the future, rather than trying to prove safety di…
x Working through a small tiling result — LessWrong Agent Foundations Löb's theorem Tiling Agents AI Frontpage 73 Working through a small tiling result by James Payor 13th May 2025 AI Alignment Forum 6 min read 9 73 Ω 38 tl;dr it seems that you can get basic tiling to work by proving that there will be safety proofs in the future, rather than trying to prove safety directly. "Tiling" here roughly refers to a state of affairs in which we have a program that is able to prove itself safe to run. I'll use a simple problem to keep this post self-contained, but here are some links to some relevant d
saved by
related reading
- Automation collapse — LessWronglesswrong.com
- Dialogue: Is there a Natural Abstraction of Good? — LessWronglesswrong.com
- Embedded Agency (full-text version) — LessWronglesswrong.com
- Definability of Truth in Probabilistic Logic - Christiano 2013intelligence.org
- Local Validity as a Key to Sanity and Civilization — LessWronglesswrong.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- As Rocks May Think | Eric Jangevjang.com
- Robust Cooperation in the Prisoner's Dilemma — LessWronglesswrong.com
- Formalizing Fermat's Last Theoremanthropic.com
- 2312.06942arxiv.org
- VingeanReflection.pdfintelligence.org
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org