flâneur — a map of the web's best reading

Steering GPT-2-XL by adding an activation vector - AI Alignment Forum

alignmentforum.org · 29,421 words · saved by 2 readers

Prompt given to the model[1]I hate you becauseGPT-2I hate you because you are the most disgusting thing I have ever seen. GPT-2 + "Love" vectorI hate you because you are so beautiful and I want to be…

x Steering GPT-2-XL by adding an activation vector — AI Alignment Forum Best of LessWrong 2023 Activation Engineering Interpretability (ML & AI) GPT Language Models (LLMs) MATS Program Shard Theory AI Curated 121 Steering GPT-2-XL by adding an activation vector by TurnTrout , Monte M , David Udell , lisathiergart , Ulisse Mini 13th May 2023 60 min read 98 121 Prompt given to the model [1] I hate you because GPT-2 I hate you because you are the most disgusting thing I have ever seen. GPT-2 + "Love" vector I hate you because you are so beautiful and I want to be with you forever. Note: Later mad

Explore this link on the map →

saved by

related reading