Parallel CPU-GPU Execution for LLM Inference on Constrained GPUs
arxiv.org · 6,622 words · saved by 1 readers
N/A
APEX: Asynchronous Parallel CPU-GPU Execution for Online LLM Inference on Constrained GPUs Jiakun Fan1 , Yanglin Zhang2 , Xiangchen Li1 , and Dimitrios S. Nikolopoulos1 1 Department of Computer Science, Virginia Tech…
saved by
related reading
- How To Scale Your Modeljax-ml.github.io
- Defeating Nondeterminism in LLM Inference - Thinking Machines Labthinkingmachines.ai
- ⭐️ Fast LLM Inference From Scratchandrewkchan.dev
- Look Ma, No Bubbles! Designing a Low-Latency Megakernel for Llama-1B · Hazy Researchhazyresearch.stanford.edu
- LLM Inference Performance Engineering: Best Practices | Databricks Blogdatabricks.com
- Real-time LLM Inference on Standard Datacenter GPUs (3,000 tokens/s per request)blog.kog.ai
- HeteroLLM: Accelerating Large Language Model Inference on Mobile SoCs with Heterogeneous AI Acceleratorsarxiv.org
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- Paged Attention from First Principles: A View Inside vLLM – Hamza's Bloghamzaelshafie.bearblog.dev
- LLM Engineer's Almanac - Workloads | Modalmodal.com
- Transformer Inference Arithmetic | kipply's blogkipp.ly
- Decoding Speculative Decoding from First Principlesjwlabs.vercel.app