flâneur — a map of the web's best reading

Mechanistically Eliciting Latent Behaviors in Language Models — LessWrong

lesswrong.com · 21,879 words · saved by 1 readers

Produced as part of the MATS Winter 2024 program, under the mentorship of Alex Turner (TurnTrout). …

x Mechanistically Eliciting Latent Behaviors in Language Models — LessWrong Activation Engineering AI Evaluations Language Models (LLMs) MATS Program AI World Modeling Frontpage 226 Mechanistically Eliciting Latent Behaviors in Language Models by Andrew Mack , TurnTrout 30th Apr 2024 AI Alignment Forum 54 min read 44 226 Ω 99 Produced as part of the MATS Winter 2024 program, under the mentorship of Alex Turner (TurnTrout). TL,DR: I introduce a method for eliciting latent behaviors in language models by learning unsupervised perturbations of an early layer of an LLM. These perturbations are tra

Explore this link on the map →

related reading