2502.13189
arxiv.org · 6,312 words · saved by 1 readers
N/A
M O BA: M IXTURE OF B LOCK ATTENTION FOR L ONG -C ONTEXT LLM S T ECHNICAL R EPORT Enzhe Lu1 Zhejun Jiang1 Jingyuan Liu1 Yulun Du1 Tao Jiang1 Chao Hong1 Shaowei Liu1 Weiran He1 Enming Yuan1 Yuzhi Wang1 Zhiqi Huang1 Huan Yuan1 arXiv:2502.13189v1 [cs.LG] 18 Feb 2025…
related reading
- [2502.13189] MoBA: Mixture of Block Attention for Long-Context LLMsarxiv.org
- The Big LLM Architecture Comparisonmagazine.sebastianraschka.com
- 2502.11089arxiv.org
- Optimizing Mixture of Block Attentionarxiv.org
- A short note on some aspects of long context attention | nor's blognor-blog.pages.dev
- Subquadratic — How SSA Makes Long Context Practicalsubq.ai
- [2603.23516] MSA: Memory Sparse Attention for Efficient End-to-End Memory Model Scaling to 100M Tokensarxiv.org
- Extending Context is Hard | kaiokendevkaiokendev.github.io
- Mamba: The Easy Wayjackcook.com
- Sparser Block-Sparse Attention via Token Permutationarxiv.org
- Transformers Inference Optimization Toolset | AstraBlogastralord.github.io
- Kimi Linear: An Expressive, Efficient Attention Architecturearxiv.org