✳flâneur — a map of the web's best reading
What’s up with LLMs representing XORs of arbitrary features? — LessWrong
lesswrong.com · 14,734 words · saved by 1 readers
Thanks to Clément Dumas, Nikola Jurković, Nora Belrose, Arthur Conmy, and Oam Patel for feedback. …
x What’s up with LLMs representing XORs of arbitrary features? — LessWrong Language Models (LLMs) AI Frontpage 159 What’s up with LLMs representing XORs of arbitrary features? by Sam Marks 3rd Jan 2024 AI Alignment Forum 19 min read 64 159 Ω 75 Thanks to Clément Dumas, Nikola Jurković, Nora Belrose, Arthur Conmy, and Oam Patel for feedback. In the comments of the post on Google Deepmind’s CCS challenges paper , I expressed skepticism that some of the experimental results seemed possible. When addressing my concerns, Rohin Shah made some claims along the lines of “If an LLM linearly represents
Explore this link on the map →related reading
- Toy Models of Superpositiontransformer-circuits.pub
- On the Biology of a Large Language Modeltransformer-circuits.pub
- Circuit Tracing: Revealing Computational Graphs in Language Modelstransformer-circuits.pub
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- What Would Non-Linear Features Actually Look Like? — Liv Gortonlivgorton.com
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- Transformer Circuits Threadtransformer-circuits.pub
- Actually, Othello-GPT Has A Linear Emergent World Representation - Neel Nandaneelnanda.io
- Structure and Interpretation of Deep Networkssidn.baulab.info
- pdfopenreview.net
- Lucius Bushnaq's Shortform — LessWronglesswrong.com
- Mapping the Mind of a Large Language Model \ Anthropicanthropic.com