✳flâneur — a map of the web's best reading
Mechanistically Eliciting Latent Behaviors in Language Models — LessWrong
lesswrong.com · 21,879 words · saved by 1 readers
Produced as part of the MATS Winter 2024 program, under the mentorship of Alex Turner (TurnTrout). …
x Mechanistically Eliciting Latent Behaviors in Language Models — LessWrong Activation Engineering AI Evaluations Language Models (LLMs) MATS Program AI World Modeling Frontpage 226 Mechanistically Eliciting Latent Behaviors in Language Models by Andrew Mack , TurnTrout 30th Apr 2024 AI Alignment Forum 54 min read 44 226 Ω 99 Produced as part of the MATS Winter 2024 program, under the mentorship of Alex Turner (TurnTrout). TL,DR: I introduce a method for eliciting latent behaviors in language models by learning unsupervised perturbations of an early layer of an LLM. These perturbations are tra
Explore this link on the map →related reading
- Mechanistically Eliciting Latent Behaviors in Language Models — AI Alignment Forumalignmentforum.org
- Emotion concepts and their function in a large language model \ Anthropicanthropic.com
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- Transformer Circuits Threadtransformer-circuits.pub
- I found >800 orthogonal “write code” steering vectors | Jacob’s Blogjacobgw.com
- Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activationstransformer-circuits.pub
- Activation Steering in 2026: A Practitioner's Field Guide | Subhadip Mitrasubhadipmitra.com
- Emotion Concepts and their Function in a Large Language Modeltransformer-circuits.pub
- Steering GPT-2-XL by adding an activation vector — AI Alignment Forumalignmentforum.org
- Eliciting Language Model Behaviors with Investigator Agents | Transluce AItransluce.org
- Surfacing Pathological Behaviors in Language Models | Transluce AItransluce.org
- Mapping the Mind of a Large Language Model \ Anthropicanthropic.com