flâneur — a map of the web's best reading

Implementing activation steering — LessWrong

lesswrong.com · 2,787 words · saved by 1 readers

Produced as part of the SERI ML Alignment Theory Scholars Program - Autumn 2023 Cohort and while being an affiliate at PIBBSS in 2024. A thank you to @Jayjay and @fela for helpful comments on this draft. This blog post is an overview of different ways to implement activation steering with some of my takes on their pros and cons. See also this GitHub repository for my minimal implementations of the different approaches. The blog post is aimed at people who are new to activation/representation steering/engineering/editing. The idea is simple: we just add some vector to the internal model activations and thus influence the model output in a similar (but sometimes more effective way) to prompting. Example[1]: Imagine that some vector in the internal representations in some transformer layer encodes a direction associated with "Love". When you add this vector to the activations of some encoded sentence "I hate the world", you change the internal representation (and thus the meaning) to some

x Implementing activation steering — LessWrong Activation Engineering Interpretability (ML & AI) Language Models (LLMs) GPT MATS Program AI Frontpage 76 Implementing activation steering by Annah 5th Feb 2024 9 min read 8 76 Produced as part of the SERI ML Alignment Theory Scholars Program - Autumn 2023 Cohort and while being an affiliate at PIBBSS in 2024. A thank you to @Jayjay and @fela for helpful comments on this draft. This blog post is an overview of different ways to implement activation steering with some of my takes on their pros and cons. See also this GitHub repository for my minima

Explore this link on the map →

saved by

related reading