TurnTrout's shortform feed — LessWrong
Comment by TurnTrout - The Scaling Monosemanticity paper doesn't do a good job comparing feature clamping to steering vectors. 1. These vectors are not "linear probes" (which are generally optimized via SGD on a logistic regression task for a supervised dataset of yes/no examples), they are difference-in-means of activation vectors 1. So call them "steering vectors"! 2. As a side note, using actual linear probe directions tends to not steer models very well (see eg Inference Time Intervention table 3 on page 8) 2. In my experience, steering vectors generally require averaging over at least 32 contrast pairs. Anthropic only compares to 1-3 contrast pairs, which is inappropriate. 1. Since feature clamping needs fewer prompts for some tasks, that is a real benefit, but you have to amortize that benefit over the huge SAE effort needed to find those features. 2. Also note that you can generate synthetic data for the steering vectors using an LLM, it isn't too hard. 3. For steering on a single task, then, steering vectors still win out in terms of amortized sample complexity (assuming the steering vectors are effective given ~32/128/256 contrast pairs, which I doubt will always be true) I totally expect feature clamping to still win out in a bunch of comparisons, it's really cool, but Anthropic's actual comparisons don't seem good and predictably underrate steering vectors. The fact that the Anthropic paper gets the comparison (and especially terminology) meaningfully wrong makes me more wary of their results going forwards.
x TurnTrout's shortform feed — LessWrong TurnTrout's shortform feed by TurnTrout 30th Jun 2019 AI Alignment Forum 1 min read 791 28 Ω 10 This is a special post for quick takes by TurnTrout . Only they can create top-level comments. Comments here also appear on the Quick Takes page and All Posts page . Rendering 0 / 791 comments, sorted by top scoring (show more) Click to highlight new comments since: Today at 8:29 AM Moderation Log More from TurnTrout View more Curated and popular this week 791 Comments 791 Comment Permalink TurnTrout 2y * Ω 33 54 13 The Scaling Monosemanticity paper doesn't d
Explore this link on the map →related reading
- Activation Steering in 2026: A Practitioner's Field Guide | Subhadip Mitrasubhadipmitra.com
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- Steering Might Stop Working Soon — LessWronglesswrong.com
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- I found >800 orthogonal “write code” steering vectors | Jacob’s Blogjacobgw.com
- Circuit Tracing: Revealing Computational Graphs in Language Modelstransformer-circuits.pub
- Steering GPT-2-XL by adding an activation vector — AI Alignment Forumalignmentforum.org
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- Dreaming Vectors: Gradient-descented steering vectors from Activation Oracles and using them to Red-Team AOs — LessWronglesswrong.com
- Steering Llama-2 with contrastive activation additions — LessWronglesswrong.com
- I found >800 orthogonal "write code" steering vectors — LessWronglesswrong.com
- What Would Non-Linear Features Actually Look Like? — Liv Gortonlivgorton.com