✳flâneur — a map of the web's best reading
Future leakage in block-quantized attention | MatX
matx.com · 1,369 words · saved by 1 readers
Quantizing attention improves efficiency on two fronts: the model has higher compute throughput, and loads fewer bytes per key/value. However, training with blo
Explore this link on the map →saved by
related reading
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- Transformer Inference Arithmetic | kipply's blogkipp.ly
- Optimizing Mixture of Block Attentionarxiv.org
- 2502.11089arxiv.org
- Overleaf Examplearxiv.org
- Transformers Inference Optimization Toolset | AstraBlogastralord.github.io
- Prompt Cache: Modular Attention Reuse for Low-Latency Inferencearxiv.org
- Linear Attention Fundamentals | Hailey Schoelkopfhaileyschoelkopf.github.io
- MiLo: Efficient Quantized MoE Inference with Mixture of Low-Rank Compensatorsbeichenhuang.github.io
- Paged Attention from First Principles: A View Inside vLLM – Hamza's Bloghamzaelshafie.bearblog.dev
- All the Transformer Math You Need to Know | How To Scale Your Modeljax-ml.github.io
- MatX: High-throughput chips for LLMsmatx.com