✳flâneur — a map of the web's best reading
The Singular Value Decompositions of Transformer Weight Matrices are Highly Interpretable — LessWrong
lesswrong.com · 15,829 words · saved by 1 readers
Please go to the colab for interactive viewing and playing with the phenomena. For space reasons, not all results included in the colab are included…
x The Singular Value Decompositions of Transformer Weight Matrices are Highly Interpretable — LessWrong Interpretability (ML & AI) Conjecture (org) AI Frontpage 200 The Singular Value Decompositions of Transformer Weight Matrices are Highly Interpretable by beren , Sid Black 28th Nov 2022 AI Alignment Forum 37 min read 34 200 Ω 69 Please go to the colab for interactive viewing and playing with the phenomena. For space reasons, not all results included in the colab are included here so please visit the colab for the full story. A GitHub repository with the colab notebook and accompanying data c
Explore this link on the map →related reading
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- Transformer Circuits Threadtransformer-circuits.pub
- The Illustrated Transformer – Jay Alammar – Visualizing machine learning one concept at a time.jalammar.github.io
- Softmax Linear Unitstransformer-circuits.pub
- Transformer Explainer: LLM Transformer Model Visually Explainedpoloclub.github.io
- Transformers from Scratche2eml.school
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- Interpreting Language Model Parametersgoodfire.ai
- Weight-Sparse Circuits May Be Interpretable Yet Unfaithful — LessWronglesswrong.com
- The Illustrated Transformer – Jay Alammar – Visualizing machine learning one concept at a time.jalammar.github.io