Mechanistically Eliciting Latent Behaviors in Language Models — AI Alignment Forum
TL,DR: I introduce a method for eliciting latent behaviors in language models by learning unsupervised perturbations of an early layer of an LLM. These perturbations are trained to maximize changes in downstream activations. The method discovers diverse and meaningful behaviors with just one prompt, including perturbations overriding safety training, eliciting backdoored behaviors and uncovering latent capabilities. Summary In the simplest case, the unsupervised perturbations I learn are given by unsupervised steering vectors - vectors added to the residual stream as a bias term in the MLP outputs of a given layer. I also report preliminary results on unsupervised steering adapters - these are LoRA adapters of the MLP output weights of a given layer, trained with the same unsupervised objective. I apply the method to several alignment-relevant toy examples, and find that the method consistently learns vectors/adapters which encode coherent and generalizable high-level behaviors. Compar
x Mechanistically Eliciting Latent Behaviors in Language Models — AI Alignment Forum Activation Engineering AI Evaluations Language Models (LLMs) MATS Program AI World Modeling Frontpage 99 Mechanistically Eliciting Latent Behaviors in Language Models by Andrew Mack , TurnTrout 30th Apr 2024 54 min read 44 99 Produced as part of the MATS Winter 2024 program, under the mentorship of Alex Turner (TurnTrout). TL,DR: I introduce a method for eliciting latent behaviors in language models by learning unsupervised perturbations of an early layer of an LLM. These perturbations are trained to maximize
saved by
related reading
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Mechanistically Eliciting Latent Behaviors in Language Models — LessWronglesswrong.com
- Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- [2501.11120] Tell me about yourself: LLMs are aware of their learned behaviorsarxiv.org
- Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activationstransformer-circuits.pub
- I found >800 orthogonal “write code” steering vectors | Jacob’s Blogjacobgw.com
- Refusal in LLMs is mediated by a single direction — LessWronglesswrong.com
- Emotion Concepts and their Function in a Large Language Modeltransformer-circuits.pub
- [2604.16812] Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- Emotion concepts and their function in a large language model \ Anthropicanthropic.com
- [2606.04071] Covert Influence Between Language Modelsarxiv.org
- Eliciting Language Model Behaviors with Investigator Agents | Transluce AItransluce.org