flâneur — a map of the web's best reading

Why K and V are not the same in Transformer attention?

stats.stackexchange.com · 1,520 words · saved by 1 readers

My understanding is for translation task K should be the same with V, but in Transformer K and V are generated by two different(randomly initialized) matrix $W^K, W^V$, therefore not the same. Can ...

neural networks - Why K and V are not the same in Transformer attention? - Cross Validated Stack Internal Knowledge at work Bring the best of human thought and AI automation together at your work. Explore Stack Internal Why K and V are not the same in Transformer attention? Ask Question Asked 6 years, 9 months ago Modified 2 years, 5 months ago Viewed 13k times 15 $\begingroup$ My understanding is for translation task K should be the same with V, but in Transformer K and V are generated by two different(randomly initialized) matrix $W^K, W^V$ , therefore not the same. Can any one tell me why?

Explore this link on the map →

related reading