2502.11089
arxiv.org · 9,470 words · saved by 3 readers
N/A
# link_1lmz0khri25.pdf ## Metadata - PDFFormatVersion=1.5 - IsLinearized=false - IsAcroFormPresent=false - IsXFAPresent=false - IsCollectionPresent=false - IsSignaturesPresent=false - CreationDate=D:20250228013606Z - Creator=LaTeX with hyperref - ModDate=D:20250228013606Z - Custom.PTEX.Fullbanner=This is pdfTeX, Version 3.141592653-2.6-1.40.25 (TeX Live 2023) kpathsea version 6.3.5 - Producer=pdfTeX-1.40.25 - Trapped=False ## Contents ### Page 1 Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse AttentionJingyang Yuan∗1,2, Huazuo Gao1, Damai Dai1, Junyu Luo2, Liang Zhao
saved by
related reading
- [2603.23516] MSA: Memory Sparse Attention for Efficient End-to-End Memory Model Scaling to 100M Tokensarxiv.org
- [2502.11089] Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attentionarxiv.org
- Overleaf Examplearxiv.org
- [2604.20920] Simplified Sparse Attention via Gist Tokensarxiv.org
- Transformers Inference Optimization Toolset | AstraBlogastralord.github.io
- [2205.14135] FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awarenessarxiv.org
- [2012.09852] SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head Pruningarxiv.org
- The Big LLM Architecture Comparisonmagazine.sebastianraschka.com
- Demystifying Sparse Attention: A Comprehensive Guide from Scratch | by VISHAL SINGH | Mediummedium.com
- Subquadratic — How SSA Makes Long Context Practicalsubq.ai
- Mamba: The Easy Wayjackcook.com
- Sparser Block-Sparse Attention via Token Permutationarxiv.org