✳flâneur — a map of the web's best reading
Overleaf Example
arxiv.org · 10,690 words · saved by 2 readers
N/A
# link_od4aaim2x1.pdf ## Metadata - PDFFormatVersion=1.5 - IsLinearized=false - IsAcroFormPresent=false - IsXFAPresent=false - IsCollectionPresent=false - IsSignaturesPresent=false - CreationDate=D:20250211015825Z - Creator=LaTeX with hyperref - ModDate=D:20250211015825Z - Custom.PTEX.Fullbanner=This is pdfTeX, Version 3.141592653-2.6-1.40.25 (TeX Live 2023) kpathsea version 6.3.5 - Producer=pdfTeX-1.40.25 - Trapped=False ## Contents ### Page 1 TREE ATTENTION: TOPOLOGY-AWARE DECODING FORLONG-CONTEXT ATTENTION ON GPU CLUSTERSVasudev Shyam1∗, Jonathan Pilault1∗, Emily Shepperd2, Quentin Antho
Explore this link on the map →saved by
related reading
- Transformer Inference Arithmetic | kipply's blogkipp.ly
- 2502.11089arxiv.org
- From Online Softmax to FlashAttentioncourses.cs.washington.edu
- Transformer Math 101 | EleutherAI Blogblog.eleuther.ai
- Transformers Inference Optimization Toolset | AstraBlogastralord.github.io
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- transformer_attention.pdfarxiv.org
- 1706.03762arxiv.org
- Linear Attention Fundamentals | Hailey Schoelkopfhaileyschoelkopf.github.io
- [2205.14135] FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awarenessarxiv.org
- [2205.14135] FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awarenessarxiv.org
- Mamba: The Easy Wayjackcook.com