flâneur — a map of the web's best reading

Future leakage in block-quantized attention | MatX

matx.com · 1,369 words · saved by 1 readers

Quantizing attention improves efficiency on two fronts: the model has higher compute throughput, and loads fewer bytes per key/value. However, training with blo

Explore this link on the map →

saved by

related reading