flâneur — a map of the web's best reading

Attribution Patching: Activation Patching At Industrial Scale — Neel Nanda

neelnanda.io · 17,147 words · saved by 6 readers

A write-up of an incomplete project I worked on at Anthropic in early 2022, using gradient-based approximation to make activation patching far more scalable

Attribution Patching: Activation Patching At Industrial Scale Feb 4 Written By Neel Nanda The following is a write-up of an (incomplete) project I worked on while at Anthropic, and a significant amount of the credit goes to the then team, Chris Olah, Catherine Olsson, Nelson Elhage & Tristan Hume. I've since cleaned up this project in my personal time and personal capacity. TLDR Activation patching is an existing technique for identifying which model activations are most important for determining model behaviour between two similar prompts that differ in a key detail I introduce a technique ca

Explore this link on the map →

saved by

related reading