flâneur — a map of the web's best reading

ceselder/dreaming-vectors: gradient-descented steering vectors from activation oracles

github.com · 528 words · saved by 1 readers

Don't get locked out of your account. Download your recovery codes or add a passkey so you don't lose access when you get a new device. There was an error while loading. Please reload this page. There was an error while loading. Please reload this page. There was an error while loading. Please reload this page. gradient-descented steering vectors from activation oracles There was an error while loading. Please reload this page. My MATS application! Full writeup on LessWrong Activation oracles (iterating on LatentQA) are an interpretability technique, capable of generating natural language explanations about activations with surprising generality. How robust are these oracles? Can we find a vector that maximises confidence of the oracle (in token probability) that a concept is represented, and can we then use this vector to steer the model? Is it possible to find feature representations that convince the oracle a concept is being represented, when it really is random noise? Eff

Dreaming Vectors: Gradient-descented steering vectors from Activation Oracles and using them to Red-Team AOs My MATS application! UPDATE: I got accepted! Full writeup on LessWrong TLDR Activation oracles (iterating on LatentQA) are an interpretability technique, capable of generating natural language explanations about activations with surprising generality. How robust are these oracles? Can we find a vector that maximises confidence of the oracle (in token probability) that a concept is represented, and can we then use this vector to steer the model? Is it possible to find feature representat

Explore this link on the map →

related reading