ceselder/dreaming-vectors: gradient-descented steering vectors from activation oracles
Don't get locked out of your account. Download your recovery codes or add a passkey so you don't lose access when you get a new device. There was an error while loading. Please reload this page. There was an error while loading. Please reload this page. There was an error while loading. Please reload this page. gradient-descented steering vectors from activation oracles There was an error while loading. Please reload this page. My MATS application! Full writeup on LessWrong Activation oracles (iterating on LatentQA) are an interpretability technique, capable of generating natural language explanations about activations with surprising generality. How robust are these oracles? Can we find a vector that maximises confidence of the oracle (in token probability) that a concept is represented, and can we then use this vector to steer the model? Is it possible to find feature representations that convince the oracle a concept is being represented, when it really is random noise? Eff
Dreaming Vectors: Gradient-descented steering vectors from Activation Oracles and using them to Red-Team AOs My MATS application! UPDATE: I got accepted! Full writeup on LessWrong TLDR Activation oracles (iterating on LatentQA) are an interpretability technique, capable of generating natural language explanations about activations with surprising generality. How robust are these oracles? Can we find a vector that maximises confidence of the oracle (in token probability) that a concept is represented, and can we then use this vector to steer the model? Is it possible to find feature representat
related reading
- Dreaming Vectors: Gradient-descented steering vectors from Activation Oracles and using them to Red-Team AOs — LessWronglesswrong.com
- Scaling Activation Oracles to Trillion-Parameter Modelstransluce.org
- [2606.02609] Building Better Activation Oraclesarxiv.org
- Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activationstransformer-circuits.pub
- Natural Language Autoencoders \ Anthropicanthropic.com
- Current activation oracles are hard to use — LessWronglesswrong.com
- Chapter 1: Transformer Interpretability - ARENAlearn.arena.education
- Emotion concepts and their function in a large language model \ Anthropicanthropic.com
- Chapter 1: Transformer Interpretability - ARENAlearn.arena.education
- [2606.02609] Building Better Activation Oraclesarxiv.org
- Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers — LessWronglesswrong.com
- I found >800 orthogonal “write code” steering vectors | Jacob’s Blogjacobgw.com