Optimizing Mixture of Block Attention
arxiv.org · 6,485 words · saved by 1 readers
N/A
Optimizing Mixture of Block Attention O PTIMIZING M IXTURE OF B LOCK ATTENTION Guangxuan Xiao1∗ Junxian Guo1∗ Kasra Mazaheri 1 Song Han1,2 1 MIT 2 NVIDIA https://github.com/mit-han-lab/flash-moba A BSTRACT Mixture of Block Attention…
related reading
- [2502.13189] MoBA: Mixture of Block Attention for Long-Context LLMsarxiv.org
- 2502.13189arxiv.org
- The Big LLM Architecture Comparisonmagazine.sebastianraschka.com
- 2502.11089arxiv.org
- Mamba: The Easy Wayjackcook.com
- Transformers Inference Optimization Toolset | AstraBlogastralord.github.io
- [2205.14135] FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awarenessarxiv.org
- LLM Inference Performance Engineering: Best Practices | Databricks Blogdatabricks.com
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- Paged Attention from First Principles: A View Inside vLLM – Hamza's Bloghamzaelshafie.bearblog.dev
- ali (@waterloo_intern) on Xx.com
- Kimi Linear: An Expressive, Efficient Attention Architecturearxiv.org