Working through a small tiling result — LessWrong
tl;dr it seems that you can get basic tiling to work by proving that there will be safety proofs in the future, rather than trying to prove safety di…
x Working through a small tiling result — LessWrong Agent Foundations Löb's theorem Tiling Agents AI Frontpage 73 Working through a small tiling result by James Payor 13th May 2025 AI Alignment Forum 6 min read 9 73 Ω 38 tl;dr it seems that you can get basic tiling to work by proving that there will be safety proofs in the future, rather than trying to prove safety directly. "Tiling" here roughly refers to a state of affairs in which we have a program that is able to prove itself safe to run. I'll use a simple problem to keep this post self-contained, but here are some links to some relevant d
Explore this link on the map →saved by
related reading
- Automation collapse — LessWronglesswrong.com
- Dialogue: Is there a Natural Abstraction of Good? — LessWronglesswrong.com
- Local Validity as a Key to Sanity and Civilization — LessWronglesswrong.com
- Critical review of Christiano's disagreements with Yudkowsky — LessWronglesswrong.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- Robust Cooperation in the Prisoner's Dilemma — LessWronglesswrong.com
- 2312.06942arxiv.org
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- Welcome!boydkane.com
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com
- An alignment safety case sketch based on debate — LessWronglesswrong.com
- 2510.01346arxiv.org