✳flâneur — a map of the web's best reading
I found >800 orthogonal "write code" steering vectors — LessWrong
lesswrong.com · 5,956 words · saved by 1 readers
Produced as part of the MATS Summer 2024 program, under the mentorship of Alex Turner (TurnTrout). …
x I found >800 orthogonal "write code" steering vectors — LessWrong Activation Engineering MATS Program AI Frontpage 114 I found >800 orthogonal "write code" steering vectors by Jacob G-W , TurnTrout 15th Jul 2024 Linkpost for jacobgw.com 9 min read 20 114 Produced as part of the MATS Summer 2024 program, under the mentorship of Alex Turner (TurnTrout). A few weeks ago, I stumbled across a very weird fact: it is possible to find multiple steering vectors in a language model that activate very similar behaviors while all being orthogonal . This was pretty surprising to me and to some people tha
Explore this link on the map →related reading
- I found >800 orthogonal “write code” steering vectors | Jacob’s Blogjacobgw.com
- Dreaming Vectors: Gradient-descented steering vectors from Activation Oracles and using them to Red-Team AOs — LessWronglesswrong.com
- Activation Steering in 2026: A Practitioner's Field Guide | Subhadip Mitrasubhadipmitra.com
- Steering GPT-2-XL by adding an activation vector — AI Alignment Forumalignmentforum.org
- Steering Might Stop Working Soon — LessWronglesswrong.com
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- A short survey on almost orthogonal vectors in a few specific large dimensionsarxiv.org
- Composer2.pdfcursor.com
- TurnTrout's shortform feed — LessWronglesswrong.com
- GitHub - ceselder/dreaming-vectors: gradient-descented steering vectors from activation oracles · GitHubgithub.com
- Transformer Circuits Threadtransformer-circuits.pub
- Emotion concepts and their function in a large language model \ Anthropicanthropic.com