Sparser Block-Sparse Attention via Token Permutation
arxiv.org · 5,188 words · saved by 1 readers
N/A
Sparser Block-Sparse Attention via Token Permutation Xinghao Wang 1 2 Pengyu Wang 1 2 Dong Zhang 1 Chenkun Tan 1 2 Shaojun Zhou 1 2 Zhaoxiang Liu 3 Shiguo Lian 3 Fangxu Liu 4 Kai Song 4 Xipeng Qiu 1 2 5 Abstract 1. Introduction Modern Large Language Models (LLMs)…
related reading
- 2502.11089arxiv.org
- [2205.14135] FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awarenessarxiv.org
- [2604.20920] Simplified Sparse Attention via Gist Tokensarxiv.org
- 2502.13189arxiv.org
- Demystifying Sparse Attention: A Comprehensive Guide from Scratch | by VISHAL SINGH | Mediummedium.com
- Overleaf Examplearxiv.org
- [2502.11089] Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attentionarxiv.org
- Subquadratic — How SSA Makes Long Context Practicalsubq.ai
- Transformers Inference Optimization Toolset | AstraBlogastralord.github.io
- The Big LLM Architecture Comparisonmagazine.sebastianraschka.com
- [2502.13189] MoBA: Mixture of Block Attention for Long-Context LLMsarxiv.org
- A short note on some aspects of long context attention | nor's blognor-blog.pages.dev