Towards a formalization of the agent structure problem — AI Alignment Forum
In Clarifying the Agent-Like Structure Problem (2022), John Wentworth describes a hypothetical instance of what he calls a selection theorem. In Scott Garrabrant's words, the question is, does agent-like behavior imply agent-like architecture? That is, if we take some class of behaving things and apply a filter for agent-like behavior, do we end up selecting things with agent-like architecture (or structure)? Of course, this question is heavily under-specified. So another way to ask this is, under which conditions does agent-like behavior imply agent-like structure? And, do those conditions feel like they formally encapsulate a naturally occurring condition? For the Q1 2024 cohort of AI Safety Camp, I was a Research Lead for a team of six people, where we worked a few hours a week to better understand and make progress on this idea. The teammates[1] were Einar Urdshals, Tyler Tracy, Jasmina Nasufi, Mateusz Bagiński, Amaury Lorin, and Alfred Harwood. The AISC project duration was too sh
x Towards a formalization of the agent structure problem — AI Alignment Forum Agent Foundations Agent-Structure Problem AI Safety Camp Optimization AI Frontpage 24 Towards a formalization of the agent structure problem by Alex_Altair 29th Apr 2024 17 min read 6 24 In Clarifying the Agent-Like Structure Problem (2022), John Wentworth describes a hypothetical instance of what he calls a selection theorem. In Scott Garrabrant 's words, the question is, does agent-like behavior imply agent-like architecture ? That is, if we take some class of behaving things and apply a filter for agent-like behav
Explore this link on the map →related reading
- Why Tool AIs Want to Be Agent AIs · Gwern.netgwern.net
- LLM Powered Autonomous Agents | Lil'Loglilianweng.github.io
- Optimality is the tiger, and agents are its teeth — LessWronglesswrong.com
- Building Effective AI Agents \ Anthropicanthropic.com
- Building reliable AI agents · parth sareenparthsareen.com
- Building Effective AI Agents \ Anthropicanthropic.com
- Agent Evaluation: A Detailed Guidecameronrwolfe.substack.com
- Beliefs are Chosen to Serve Goals — LessWronglesswrong.com
- [1902.09469] Embedded Agencyarxiv.org
- Arjun Virkarjunvirk.com
- Embedded Agency (full-text version) — LessWronglesswrong.com
- A Taxonomy of RL Environments for LLM Agentsleehanchung.github.io