✳flâneur — a map of the web's best reading
2502.11089
arxiv.org · 9,470 words · saved by 3 readers
N/A
# link_1lmz0khri25.pdf ## Metadata - PDFFormatVersion=1.5 - IsLinearized=false - IsAcroFormPresent=false - IsXFAPresent=false - IsCollectionPresent=false - IsSignaturesPresent=false - CreationDate=D:20250228013606Z - Creator=LaTeX with hyperref - ModDate=D:20250228013606Z - Custom.PTEX.Fullbanner=This is pdfTeX, Version 3.141592653-2.6-1.40.25 (TeX Live 2023) kpathsea version 6.3.5 - Producer=pdfTeX-1.40.25 - Trapped=False ## Contents ### Page 1 Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse AttentionJingyang Yuan∗1,2, Huazuo Gao1, Damai Dai1, Junyu Luo2, Liang Zhao
Explore this link on the map →saved by
related reading
- Overleaf Examplearxiv.org
- [2502.11089] Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attentionarxiv.org
- Transformers Inference Optimization Toolset | AstraBlogastralord.github.io
- [2012.09852] SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head Pruningarxiv.org
- The Bitter Lesson is coming for Tokenization – ⛰️ lucalplucalp.dev
- Demystifying Sparse Attention: A Comprehensive Guide from Scratch | by VISHAL SINGH | Mediummedium.com
- Subquadratic — How SSA Makes Long Context Practicalsubq.ai
- On the Tradeoffs of SSMs and Transformers | Goomba Labgoombalab.github.io
- Mamba: The Easy Wayjackcook.com
- Sparse Attention Post-Training for Mechanistic Interpretabilityarxiv.org
- Linear Attention Fundamentals | Hailey Schoelkopfhaileyschoelkopf.github.io
- Pengyu Zhao on X: "MiniMax M2 Tech Blog 3: Why Did M2 End Up as a Full Attention Model? On behave of pre-training lead Haohai Sun. (https://t.co/WH4xOD9KrT) I. Introduction As the lead of MiniMax-M2 pretrain, I've been getting many queries x.com