flâneur — a map of the web's best reading

Simplex Progress Report - July 2025 — LessWrong

lesswrong.com · 5,357 words · saved by 1 readers

At Simplex our mission is to develop a principled science of the representations and emergent behaviors of AI systems. Our initial work showed that transformers linearly represent belief state geometries in their residual streams. We think of that work as providing the first steps into an understanding of what fundamentally we are training AI systems to do, and what representations we are training them to have. Since that time, we have used that framework to make progress in a number of directions, which we will present in the sections below. The projects ask, and provide answers to, the following questions: In answering these questions we find a number of surprising results, described in the rest of this post, and new ways of thinking about both specific parts of the transformer architecture, for instance what exactly the attention heads are doing, and more generally computation in neural networks overall that we believe have important implications for much of interpretability. We are

x Simplex Progress Report - July 2025 — LessWrong Interpretability (ML & AI) AI Personal Blog 2025 Top Fifty: 14 % 116 Simplex Progress Report - July 2025 by Adam Shai , Paul Riechers , hrbigelow , Eric Alt , mntss 28th Jul 2025 AI Alignment Forum 18 min read 3 116 Ω 49 Thanks to Jasmina Urdshals, Xavier Poncini, and Justis Mills for comments. Introduction At Simplex our mission is to develop a principled science of the representations and emergent behaviors of AI systems. Our initial work showed that transformers linearly represent belief state geometries in their residual stream s. We think

Explore this link on the map →

related reading