A Telescope, Pointed at the Stars | Tilde
We believe mechanistic understanding is the foundation for entirely new architectures, capabilities, and safety tools. Endlessly throwing tricks and compute at problems can get you to the next benchmark, but it rarely tells you why something works or how to build on it. The most powerful breakthroughs will come from deeply understanding how models work. The past year has witnessed the evolution of models at breathtaking speed. The strongest example would perhaps be the emergence and ongoing development of reasoning models, with the initial release of OpenAI’s o1 and the subsequent proliferation of test-time scaling approaches (Deepseek R1, Qwen 3, S1, etc.)[1][2][3][4]. In the world of open-source, we finally have open-weight MoE models at frontier scale/quality, released by Qwen, Kimi, Deepseek, and recently OpenAI as well[5][6]. It’s an exciting time to be studying the inner workings of models, with such strong models freely available. The past 2 years have also seen amazing progress
Back 8.10.2025 Ben Keigwin, Dhruv Pai Thesis We believe mechanistic understanding is the foundation for entirely new architectures, capabilities, and safety tools. Endlessly throwing tricks and compute at problems can get you to the next benchmark, but it rarely tells you why something works or how to build on it. The most powerful breakthroughs will come from deeply understanding how models work. Acceleration The past year has witnessed the evolution of models at breathtaking speed. The strongest example would perhaps be the emergence and ongoing development of reasoning models, with…
saved by
related reading
- Transformer Circuits Threadtransformer-circuits.pub
- Dario Amodei — The Urgency of Interpretabilitydarioamodei.com
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- How To Become A Mechanistic Interpretability Researcher — AI Alignment Forumalignmentforum.org
- As Rocks May Think | Eric Jangevjang.com
- A Comprehensive Mechanistic Interpretability Explainer & Glossary — Neel Nandaneelnanda.io
- A Pragmatic Vision for Interpretability — LessWronglesswrong.com
- On Optimism for Interpretabilitygoodfire.ai
- An Ambitious Vision for Interpretability — AI Alignment Forumalignmentforum.org
- Neel Nanda on Mechanistic Interpretability: Progress, Limits, and Paths to Safer AI — EA Forumforum.effectivealtruism.org