Mechanistically Eliciting Latent Behaviors in Language Models — AI Alignment Forum
TL,DR: I introduce a method for eliciting latent behaviors in language models by learning unsupervised perturbations of an early layer of an LLM. These perturbations are trained to maximize changes in downstream activations. The method discovers diverse and meaningful behaviors with just one prompt, including perturbations overriding safety training, eliciting backdoored behaviors and uncovering latent capabilities. Summary In the simplest case, the unsupervised perturbations I learn are given by unsupervised steering vectors - vectors added to the residual stream as a bias term in the MLP outputs of a given layer. I also report preliminary results on unsupervised steering adapters - these are LoRA adapters of the MLP output weights of a given layer, trained with the same unsupervised objective. I apply the method to several alignment-relevant toy examples, and find that the method consistently learns vectors/adapters which encode coherent and generalizable high-level behaviors. Compar
x Mechanistically Eliciting Latent Behaviors in Language Models — AI Alignment Forum Activation Engineering AI Evaluations Language Models (LLMs) MATS Program AI World Modeling Frontpage 99 Mechanistically Eliciting Latent Behaviors in Language Models by Andrew Mack , TurnTrout 30th Apr 2024 54 min read 44 99 Produced as part of the MATS Winter 2024 program, under the mentorship of Alex Turner (TurnTrout). TL,DR: I introduce a method for eliciting latent behaviors in language models by learning unsupervised perturbations of an early layer of an LLM. These perturbations are trained to maximize
Explore this link on the map →related reading
- Mechanistically Eliciting Latent Behaviors in Language Models — LessWronglesswrong.com
- Emotion concepts and their function in a large language model \ Anthropicanthropic.com
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- Transformer Circuits Threadtransformer-circuits.pub
- Eliciting Language Model Behaviors with Investigator Agents | Transluce AItransluce.org
- Surfacing Pathological Behaviors in Language Models | Transluce AItransluce.org
- Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activationstransformer-circuits.pub
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Activation Steering in 2026: A Practitioner's Field Guide | Subhadip Mitrasubhadipmitra.com
- Emotion Concepts and their Function in a Large Language Modeltransformer-circuits.pub
- Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- I found >800 orthogonal “write code” steering vectors | Jacob’s Blogjacobgw.com