AlgZoo: uninterpreted models with fewer than 1,500 parameters — LessWrong
This post covers work done by several researchers at, visitors to and collaborators of ARC, including Zihao Chen, George Robinson, David Matolcsi, Ja…
x AlgZoo: uninterpreted models with fewer than 1,500 parameters — LessWrong Interpretability (ML & AI) Alignment Research Center (ARC) AI Frontpage 2026 Top Fifty: 14 % 181 AlgZoo: uninterpreted models with fewer than 1,500 parameters by Jacob_Hilton 26th Jan 2026 ARC AI Alignment Forum 12 min read 7 181 Ω 61 This post covers work done by several researchers at, visitors to and collaborators of ARC, including Zihao Chen, George Robinson, David Matolcsi, Jacob Stavrianos, Jiawei Li and Michael Sklar. Thanks to Aryan Bhatt, Gabriel Wu, Jiawei Li, Lee Sharkey, Victor Lecomte and Zihao Chen for co
Explore this link on the map →saved by
related reading
- ARC progress update: Competing with sampling — LessWronglesswrong.com
- Formal verification, heuristic explanations and surprise accounting — Alignment Research Centeralignment.org
- Attribution-based parameter decomposition — LessWronglesswrong.com
- Competing with sampling — Alignment Research Centeralignment.org
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- The Unreasonable Effectiveness of Recurrent Neural Networkskarpathy.github.io
- Transformer Circuits Threadtransformer-circuits.pub
- Aman's AI Journal • Primers • Ilya Sutskever's Top 30aman.ai
- Interpreting Language Model Parametersgoodfire.ai
- the-illusion-of-thinking.pdfml-site.cdn-apple.com
- Language models can explain neurons in language modelsopenaipublic.blob.core.windows.net