flâneur — a map of the web's best reading

Steering GPT-2-XL by adding an activation vector — LessWrong

lesswrong.com · 25,311 words · saved by 1 readers

Alex Turner and collaborators show that you can modify GPT-2's behavior in surprising and interesting ways by just adding activation vectors to its f…

x Steering GPT-2-XL by adding an activation vector — LessWrong Best of LessWrong 2023 Activation Engineering Interpretability (ML & AI) GPT Language Models (LLMs) MATS Program Shard Theory AI Curated 442 Steering GPT-2-XL by adding an activation vector by TurnTrout , Monte M , David Udell , lisathiergart , Ulisse Mini 13th May 2023 AI Alignment Forum 60 min read 98 442 Ω 121 Prompt given to the model [1] I hate you because GPT-2 I hate you because you are the most disgusting thing I have ever seen. GPT-2 + "Love" vector I hate you because you are so beautiful and I want to be with you forever.

Explore this link on the map →

related reading