A bird's eye view of ARC's research — Alignment Research Center
Over the last few months, ARC has released a number of pieces of research. While some of these can be independently motivated, there is also a more unified research vision behind them. The purpose of this post is to try to convey some of that vision and how our individual pieces of research fit into it. Thanks to Ryan Greenblatt, Victor Lecomte, Eric Neyman, Jeff Wu and Mark Xu for helpful comments. To begin, we will take a "bird's eye" view of ARC's research using an interactive diagram. 1 As you "zoom in", more nodes will become visible and the explanation below will update to explain the new nodes. At the most zoomed-out level, ARC is working on the problem of "intent alignment": how to design AI systems that are trying to do what their operators want. While many practitioners are taking an iterative approach to this problem, there are foreseeable ways in which today's leading approaches could fail to scale to more intelligent AI systems, which could have undesirable consequences.
Over the last few months, ARC has released a number of pieces of research. While some of these can be independently motivated, there is also a more unified research vision behind them. The purpose of this post is to try to convey some of that vision and how our individual pieces of research fit into it. Thanks to Ryan Greenblatt, Victor Lecomte, Eric Neyman, Jeff Wu and Mark Xu for helpful comments. A bird's eye view To begin, we will take a "bird's eye" view of ARC's research using an interactive diagram. [1] As you "zoom in", more nodes will become visible and the explanation below will upda
Explore this link on the map →related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Obstacles in ARC's agenda: Finding explanations — LessWronglesswrong.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- A Mike's-Eye View of ARC's Research — Alignment Research Centeralignment.org
- Mediumai-alignment.com
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- Mechanistic anomaly detection and ELK — LessWronglesswrong.com
- ARC progress update: Competing with sampling — LessWronglesswrong.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- (My understanding of) What Everyone in Technical Alignment is Doing and Why — LessWronglesswrong.com
- Competing with sampling — Alignment Research Centeralignment.org
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org