Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference
arxiv.org · 5,191 words · saved by 1 readers
N/A
Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference Jiaming Tang * 1 2 Yilong Zhao * 1 3 Kan Zhu 3 Guangxuan Xiao 2 Baris Kasikci 3 Song Han 2 4 Abstract As the demand for long-context large language models (LLMs) increases, models with context arXiv:2406.10774v2 [cs.CL] 26 Aug 2024 windows of up to 128K or 1M tokens…
related reading
- 2502.11089arxiv.org
- Paged Attention from First Principles: A View Inside vLLM – Hamza's Bloghamzaelshafie.bearblog.dev
- Subquadratic — How SSA Makes Long Context Practicalsubq.ai
- [2412.02252] Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarityarxiv.org
- [2506.06266] Cartridges: Lightweight and general-purpose long context representations via self-studyarxiv.org
- [2604.20920] Simplified Sparse Attention via Gist Tokensarxiv.org
- 2407.15891arxiv.org
- LLM Inference Performance Engineering: Best Practices | Databricks Blogdatabricks.com
- [2412.14838] DynamicKV: Task-Aware Adaptive KV Cache Compression for Long Context LLMsarxiv.org
- Transformers Inference Optimization Toolset | AstraBlogastralord.github.io
- Star Attention: Efficient LLM Inference over Long Sequencesarxiv.org
- [2603.23516] MSA: Memory Sparse Attention for Efficient End-to-End Memory Model Scaling to 100M Tokensarxiv.org