Circuits Updates - April 2025
We report a number of developing ideas on the Anthropic interpretability team, which might be of interest to researchers working actively in this space. Some of these are emerging strands of research where we expect to publish more on in the coming months. Others are minor points we wish to share, since we're unlikely to ever write a paper about them. We'd ask you to treat these results like those of a colleague sharing some thoughts or preliminary experiments for a few minutes at a lab meeting, rather than a mature paper. New Posts In “On the Biology of a Large Language Model”, we examined a jailbreak in which Haiku partially complies with a request to explain how to make a bomb, and compared it to a baseline prompt in which Haiku is directly asked for instructions for making a bomb, but refuses. One prompt we examined but did not include in the paper is an unsuccessful jailbreak which we discuss below. Surprisingly, by applying the same circuit tracing methodology to this prompt, we
Circuits Updates - April 2025 Transformer Circuits Thread Circuits Updates - April 2025 We report a number of developing ideas on the Anthropic interpretability team, which might be of interest to researchers working actively in this space. Some of these are emerging strands of research where we expect to publish more on in the coming months. Others are minor points we wish to share, since we're unlikely to ever write a paper about them. We'd ask you to treat these results like those of a colleague sharing some thoughts or preliminary experiments for a few minutes at a lab meeting, rather than
Explore this link on the map →related reading
- Transformer Circuits Threadtransformer-circuits.pub
- On the Biology of a Large Language Modeltransformer-circuits.pub
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- Circuits Updates - July 2024transformer-circuits.pub
- Circuit Tracing: Revealing Computational Graphs in Language Modelstransformer-circuits.pub
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- Dario Amodei — The Urgency of Interpretabilitydarioamodei.com
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Mapping the Mind of a Large Language Model \ Anthropicanthropic.com
- Circuits Updates - April 2024transformer-circuits.pub
- Neel Nanda on Mechanistic Interpretability: Progress, Limits, and Paths to Safer AI — EA Forumforum.effectivealtruism.org