Context Parallelism for Scalable Million-Token Inference
arxiv.org · 5,432 words · saved by 1 readers
N/A
C ONTEXT PARALLELISM FOR S CALABLE M ILLION -T OKEN I NFERENCE Amy (Jie) Yang 1 Jingyi Yang 1 Aya Ibrahim 1 Xinfeng Xie 1 Bangsheng Tang 1 Grigory Sizov 1 Jeremy Reizenstein 1 Jongsoo Park 1 Jianyu Huang 1 A BSTRACT We present context parallelism for long-context large language model inference, which achieves near-linear…
related reading
- Transformer Inference Arithmetic | kipply's blogkipp.ly
- Defeating Nondeterminism in LLM Inference - Thinking Machines Labthinkingmachines.ai
- TokenRing: An Efficient Parallelism Framework for Infinite-Context LLMs via Bidirectional Communicationarxiv.org
- Paged Attention from First Principles: A View Inside vLLM – Hamza's Bloghamzaelshafie.bearblog.dev
- LLM Inference Performance Engineering: Best Practices | Databricks Blogdatabricks.com
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- [2412.02252] Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarityarxiv.org
- ⭐️ Fast LLM Inference From Scratchandrewkchan.dev
- 2310.01889arxiv.org
- Mediumblog.gopenai.com
- Overleaf Examplearxiv.org
- Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inferencearxiv.org