2407.15891
arxiv.org · 6,403 words · saved by 1 readers
N/A
R AZOR ATTENTION : E FFICIENT KV C ACHE C OMPRESSION T HROUGH R ETRIEVAL H EADS A P REPRINT Hanlin Tang* 1 , Yang Lin1 , Jing Lin1 , Qingsen Han1 , Shikuan Hong1 , Yiwu Yao1 , and Gongyi Wang1 1…
related reading
- [2412.02252] Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarityarxiv.org
- TurboQuant: Redefining AI efficiency with extreme compressionresearch.google
- [2412.14838] DynamicKV: Task-Aware Adaptive KV Cache Compression for Long Context LLMsarxiv.org
- 2502.11089arxiv.org
- Paged Attention from First Principles: A View Inside vLLM – Hamza's Bloghamzaelshafie.bearblog.dev
- The Big LLM Architecture Comparisonmagazine.sebastianraschka.com
- Transformers Inference Optimization Toolset | AstraBlogastralord.github.io
- ali (@waterloo_intern) on Xx.com
- Cache strategies · Hugging Facehuggingface.co
- TurboQuant: Redefining AI efficiency with extreme compressionresearch.google
- 2306.14048arxiv.org
- KV Caching Explained: Optimizing Transformer Inference Efficiencyhuggingface.co