flâneur — a map of the web's best reading

Towards a formalization of the agent structure problem — AI Alignment Forum

alignmentforum.org · 4,617 words · saved by 1 readers

In Clarifying the Agent-Like Structure Problem (2022), John Wentworth describes a hypothetical instance of what he calls a selection theorem. In Scott Garrabrant's words, the question is, does agent-like behavior imply agent-like architecture? That is, if we take some class of behaving things and apply a filter for agent-like behavior, do we end up selecting things with agent-like architecture (or structure)? Of course, this question is heavily under-specified. So another way to ask this is, under which conditions does agent-like behavior imply agent-like structure? And, do those conditions feel like they formally encapsulate a naturally occurring condition? For the Q1 2024 cohort of AI Safety Camp, I was a Research Lead for a team of six people, where we worked a few hours a week to better understand and make progress on this idea. The teammates[1] were Einar Urdshals, Tyler Tracy, Jasmina Nasufi, Mateusz Bagiński, Amaury Lorin, and Alfred Harwood. The AISC project duration was too sh

x Towards a formalization of the agent structure problem — AI Alignment Forum Agent Foundations Agent-Structure Problem AI Safety Camp Optimization AI Frontpage 24 Towards a formalization of the agent structure problem by Alex_Altair 29th Apr 2024 17 min read 6 24 In Clarifying the Agent-Like Structure Problem (2022), John Wentworth describes a hypothetical instance of what he calls a selection theorem. In Scott Garrabrant 's words, the question is, does agent-like behavior imply agent-like architecture ? That is, if we take some class of behaving things and apply a filter for agent-like behav

Explore this link on the map →

related reading