HeteroLLM: Accelerating Large Language Model Inference on Mobile SoCs with Heterogeneous AI Accelerators
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions. HTML conversions sometimes display errors due to content that did not convert correctly from the source. This paper uses the following packages that are not yet supported by the HTML conversion tool. Feedback on these issues are not necessary; they are known and are being worked on. Authors: achieve the best HTML results from your LaTeX submissions by following these best practices. With the rapid advancement of artificial intelligence technologies such as ChatGPT, AI agents and video generation, contemporary mobile systems have begun integrating these AI capabilities on local devices to enhance privacy and reduce response latency. To meet the computational demands of AI tasks, current mobile SoCs are equipped with diverse AI accelerators, including GPUs
HeteroLLM: Accelerating Large Language Model Inference on Mobile SoCs with Heterogeneous AI Accelerators Le Chen 1† , Dahu Feng 1§ , Erhu Feng † , Rong Zhao § , Yingrui Wang ‡ Yubin Xia ∗† , Haibo Chen † , Pinjie Xu ⋄‡ † Institute of Parallel and Distributed Systems, Shanghai Jiao Tong University § Tsinghua University ‡ SenseTime Research Abstract With the rapid advancement of artificial intelligence technologies such as ChatGPT, AI agents and video generation, contemporary mobile systems have begun integrating these AI capabilities on local devices to enhance privacy and reduce response laten
Explore this link on the map →saved by
related reading
- ⭐️ Fast LLM Inference From Scratchandrewkchan.dev
- How To Scale Your Modeljax-ml.github.io
- Defeating Nondeterminism in LLM Inference - Thinking Machines Labthinkingmachines.ai
- [2403.00579] NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM Inferencingarxiv.org
- LLM Inference Performance Engineering: Best Practices | Databricks Blogdatabricks.com
- Transformer Inference Arithmetic | kipply's blogkipp.ly
- Look Ma, No Bubbles! Designing a Low-Latency Megakernel for Llama-1B · Hazy Researchhazyresearch.stanford.edu
- How is LLaMa.cpp possible?finbarr.ca
- Real-time LLM Inference on Standard Datacenter GPUs (3,000 tokens/s per request)blog.kog.ai
- How to Think About GPUs | How To Scale Your Modeljax-ml.github.io
- Unlocking the full power of NVIDIA H100 GPUs for ML inference with TensorRTbaseten.co
- Best practices to accelerate inference for large-scale production workloadstogether.ai