flâneur — a map of the web's best reading

Circuits Updates - April 2025

transformer-circuits.pub · 3,774 words · saved by 1 readers

We report a number of developing ideas on the Anthropic interpretability team, which might be of interest to researchers working actively in this space. Some of these are emerging strands of research where we expect to publish more on in the coming months. Others are minor points we wish to share, since we're unlikely to ever write a paper about them. We'd ask you to treat these results like those of a colleague sharing some thoughts or preliminary experiments for a few minutes at a lab meeting, rather than a mature paper. New Posts In “On the Biology of a Large Language Model”, we examined a jailbreak in which Haiku partially complies with a request to explain how to make a bomb, and compared it to a baseline prompt in which Haiku is directly asked for instructions for making a bomb, but refuses. One prompt we examined but did not include in the paper is an unsuccessful jailbreak which we discuss below. Surprisingly, by applying the same circuit tracing methodology to this prompt, we

Circuits Updates - April 2025 Transformer Circuits Thread Circuits Updates - April 2025 We report a number of developing ideas on the Anthropic interpretability team, which might be of interest to researchers working actively in this space. Some of these are emerging strands of research where we expect to publish more on in the coming months. Others are minor points we wish to share, since we're unlikely to ever write a paper about them. We'd ask you to treat these results like those of a colleague sharing some thoughts or preliminary experiments for a few minutes at a lab meeting, rather than

Explore this link on the map →

related reading