Competing with sampling — Alignment Research Center
In 2025, ARC has been making conceptual and theoretical progress at the fastest pace that I've seen since I first interned in 2022. Most of this progress has come about because of a re-orientation around a more specific goal: outperforming random sampling when it comes to understanding neural
In 2025, ARC has been making conceptual and theoretical progress at the fastest pace that I've seen since I first interned in 2022. Most of this progress has come about because of a re-orientation around a more specific goal: outperforming random sampling when it comes to understanding neural network outputs. Compared to our previous goals, this goal has the advantage of being more concrete and more directly tied to useful applications. The purpose of this post is to: Explain and motivate our "outperforming sampling" agenda from the standpoint of preventing catastrophic AI misalignment. Introd
saved by
related reading
- ARC progress update: Competing with sampling — LessWronglesswrong.com
- A Mike's-Eye View of ARC's Research — Alignment Research Centeralignment.org
- AlgZoo: uninterpreted models with fewer than 1,500 parameters — LessWronglesswrong.com
- [2605.05179] Estimating the expected output of wide random MLPs more efficiently than samplingarxiv.org
- Announcing the ARC White-Box Estimation Challenge — LessWronglesswrong.com
- An optimization perspective on log-concave sampling and beyond | Sinho Chewichewisinho.github.io
- Alignment Research Centeralignment.org
- As Rocks May Think | Eric Jangevjang.com
- Fact Finding: Attempting to Reverse-Engineer Factual Recall on the Neuron Level (Post 1) — AI Alignment Forumalignmentforum.org
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- arxiv.org/pdf/1805.08522arxiv.org
- Research update: Towards a Law of Iterated Expectations for Heuristic Estimators — Alignment Research Centeralignment.org