Star Attention: Efficient LLM Inference over Long Sequences
arxiv.org · 4,505 words · saved by 1 readers
N/A
Star Attention: Efficient LLM Inference over Long Sequences Shantanu Acharya 1 Fei Jia 1 Boris Ginsburg 1 Abstract these methods improve training efficiency, autoregressive decoding during inference still requires the model to attend Inference with…
related reading
- Transformers Inference Optimization Toolset | AstraBlogastralord.github.io
- 2502.11089arxiv.org
- [2205.14135] FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awarenessarxiv.org
- Paged Attention from First Principles: A View Inside vLLM – Hamza's Bloghamzaelshafie.bearblog.dev
- LLM Inference Performance Engineering: Best Practices | Databricks Blogdatabricks.com
- Overleaf Examplearxiv.org
- Getting Caught Up to Modern LLM Research | Samarth Goeldev.samarthgoel.com
- ⭐️ Fast LLM Inference From Scratchandrewkchan.dev
- Subquadratic — How SSA Makes Long Context Practicalsubq.ai
- [2604.20920] Simplified Sparse Attention via Gist Tokensarxiv.org
- How LLM Inference Worksarpitbhayani.me
- Log-Linear Attentionarxiv.org