Why K and V are not the same in Transformer attention?
My understanding is for translation task K should be the same with V, but in Transformer K and V are generated by two different(randomly initialized) matrix $W^K, W^V$, therefore not the same. Can ...
neural networks - Why K and V are not the same in Transformer attention? - Cross Validated Stack Internal Knowledge at work Bring the best of human thought and AI automation together at your work. Explore Stack Internal Why K and V are not the same in Transformer attention? Ask Question Asked 6 years, 9 months ago Modified 2 years, 5 months ago Viewed 13k times 15 $\begingroup$ My understanding is for translation task K should be the same with V, but in Transformer K and V are generated by two different(randomly initialized) matrix $W^K, W^V$ , therefore not the same. Can any one tell me why?
Explore this link on the map →related reading
- neural networks - What exactly are keys, queries, and values in attention mechanisms? - Cross Validatedstats.stackexchange.com
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- The Illustrated Transformer – Jay Alammar – Visualizing machine learning one concept at a time.jalammar.github.io
- Transformers from Scratche2eml.school
- The Illustrated Transformer – Jay Alammar – Visualizing machine learning one concept at a time.jalammar.github.io
- transformer_attention.pdfarxiv.org
- Multi Query Attention (MQA) and Grouped-Query Attention (GQA)tinkerd.net
- A Conceptual Guide to Transformers: Part Ibenlevinstein.substack.com
- An Intuition for Attention | Jay Modyjaykmody.com
- Transformers Explained Visually (Part 3): Multi-head Attention, deep dive | Towards Data Sciencetowardsdatascience.com
- 1706.03762arxiv.org
- Everything About Transformerskrupadave.com