Simplex Progress Report - July 2025 — LessWrong
At Simplex our mission is to develop a principled science of the representations and emergent behaviors of AI systems. Our initial work showed that transformers linearly represent belief state geometries in their residual streams. We think of that work as providing the first steps into an understanding of what fundamentally we are training AI systems to do, and what representations we are training them to have. Since that time, we have used that framework to make progress in a number of directions, which we will present in the sections below. The projects ask, and provide answers to, the following questions: In answering these questions we find a number of surprising results, described in the rest of this post, and new ways of thinking about both specific parts of the transformer architecture, for instance what exactly the attention heads are doing, and more generally computation in neural networks overall that we believe have important implications for much of interpretability. We are
x Simplex Progress Report - July 2025 — LessWrong Interpretability (ML & AI) AI Personal Blog 2025 Top Fifty: 14 % 116 Simplex Progress Report - July 2025 by Adam Shai , Paul Riechers , hrbigelow , Eric Alt , mntss 28th Jul 2025 AI Alignment Forum 18 min read 3 116 Ω 49 Thanks to Jasmina Urdshals, Xavier Poncini, and Justis Mills for comments. Introduction At Simplex our mission is to develop a principled science of the representations and emergent behaviors of AI systems. Our initial work showed that transformers linearly represent belief state geometries in their residual stream s. We think
Explore this link on the map →related reading
- Transformers Represent Belief State Geometry in their Residual Stream — LessWronglesswrong.com
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- Transformer Circuits Threadtransformer-circuits.pub
- On the Tradeoffs of SSMs and Transformers | Goomba Labgoombalab.github.io
- Transformers from Scratche2eml.school
- Transformer Explainer: LLM Transformer Model Visually Explainedpoloclub.github.io
- In-context Learning and Induction Headstransformer-circuits.pub
- The Topological Trouble With Transformersarxiv.org
- Simplex - building our research team | Manifundmanifund.org
- Understanding Transformers... (beyond the Math) – kalomaze's kalomazing blogkalomaze.bearblog.dev
- The World Inside Neural Networksgoodfire.ai
- GenAI Handbookgenai-handbook.github.io