Obstacles in ARC's agenda: Finding explanations — LessWrong
As an employee of the European AI Office, it's important for me to emphasize this point: The views and opinions of the author expressed herein are pe…
x Obstacles in ARC's agenda: Finding explanations — LessWrong Obstacles in ARC's agenda Alignment Research Center (ARC) Eliciting Latent Knowledge AI Frontpage 2025 Top Fifty: 14 % 128 Obstacles in ARC's agenda: Finding explanations by David Matolcsi 30th Apr 2025 20 min read 10 128 As an employee of the European AI Office, it's important for me to emphasize this point: The views and opinions of the author expressed herein are personal and do not necessarily reflect those of the European Commission or other EU institutions. Also, to stave off a common confusion: I worked at ARC Theory, which i
Explore this link on the map →related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- A bird's eye view of ARC's research — Alignment Research Centeralignment.org
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- Formal verification, heuristic explanations and surprise accounting — Alignment Research Centeralignment.org
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- ARC progress update: Competing with sampling — LessWronglesswrong.com
- An Ambitious Vision for Interpretability — AI Alignment Forumalignmentforum.org
- Competing with sampling — Alignment Research Centeralignment.org
- A Mike's-Eye View of ARC's Research — Alignment Research Centeralignment.org
- How Can Interpretability Researchers Help AGI Go Well? — AI Alignment Forumalignmentforum.org
- Neel Nanda on Mechanistic Interpretability: Progress, Limits, and Paths to Safer AI — EA Forumforum.effectivealtruism.org
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com