✳flâneur — a map of the web's best reading
Transformer Circuits Thread
transformer-circuits.pub · 1,382 words · saved by 14 readers
Can we reverse engineer transformer language models into human-understandable computer programs?
Transformer Circuits Thread --> Transformer Circuits Thread Anthropic’s Interpretability Research A surprising fact about modern large language models is that nobody really knows how they work internally. The Interpretability team strives to change that — to understand these models to better plan for a future of safe AI. May 2026 Circuits Updates — May 2026 A short update on understanding features through downstream connections. Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations Fraser-Taliente, Kantamneni, Ong, et al., 2026 We train Claude to translate its inte
Explore this link on the map →saved by
- Bryan Chiang
- Shubham Shah
- Kaylee George
- Karan MJ
- Claire Wang
- Asma Lamgh
- Emma Guo
- Asher P
- Jordan Sucher
- Lydia Nottingham
- Jirat C
- Nathan Chen
related reading
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- On the Biology of a Large Language Modeltransformer-circuits.pub
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Circuit Tracing: Revealing Computational Graphs in Language Modelstransformer-circuits.pub
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- In-context Learning and Induction Headstransformer-circuits.pub
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Mapping the Mind of a Large Language Model \ Anthropicanthropic.com
- Circuits Updates - January 2024transformer-circuits.pub
- Transformer Explainer: LLM Transformer Model Visually Explainedpoloclub.github.io