ceselder/dreaming-vectors: gradient-descented steering vectors from activation oracles
Don't get locked out of your account. Download your recovery codes or add a passkey so you don't lose access when you get a new device. There was an error while loading. Please reload this page. There was an error while loading. Please reload this page. There was an error while loading. Please reload this page. gradient-descented steering vectors from activation oracles There was an error while loading. Please reload this page. My MATS application! Full writeup on LessWrong Activation oracles (iterating on LatentQA) are an interpretability technique, capable of generating natural language explanations about activations with surprising generality. How robust are these oracles? Can we find a vector that maximises confidence of the oracle (in token probability) that a concept is represented, and can we then use this vector to steer the model? Is it possible to find feature representations that convince the oracle a concept is being represented, when it really is random noise? Eff
Dreaming Vectors: Gradient-descented steering vectors from Activation Oracles and using them to Red-Team AOs My MATS application! UPDATE: I got accepted! Full writeup on LessWrong TLDR Activation oracles (iterating on LatentQA) are an interpretability technique, capable of generating natural language explanations about activations with surprising generality. How robust are these oracles? Can we find a vector that maximises confidence of the oracle (in token probability) that a concept is represented, and can we then use this vector to steer the model? Is it possible to find feature representat
Explore this link on the map →related reading
- Dreaming Vectors: Gradient-descented steering vectors from Activation Oracles and using them to Red-Team AOs — LessWronglesswrong.com
- Natural Language Autoencoders \ Anthropicanthropic.com
- Current activation oracles are hard to use — LessWronglesswrong.com
- Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activationstransformer-circuits.pub
- Chapter 1: Transformer Interpretability - ARENAlearn.arena.education
- Emotion concepts and their function in a large language model \ Anthropicanthropic.com
- Chapter 1: Transformer Interpretability - ARENAlearn.arena.education
- Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers — LessWronglesswrong.com
- I found >800 orthogonal “write code” steering vectors | Jacob’s Blogjacobgw.com
- Activation Steering in 2026: A Practitioner's Field Guide | Subhadip Mitrasubhadipmitra.com
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- [2606.02609] Building Better Activation Oraclesarxiv.org