[2403.00579] NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM Inferencing
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs. arXiv Operational Status Get status notifications via email or slack
[2403.00579] NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM Inferencing --> Computer Science > Hardware Architecture arXiv:2403.00579 (cs) [Submitted on 1 Mar 2024 ( v1 ), last revised 29 Mar 2024 (this version, v3)] Title: NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM Inferencing Authors: Guseul Heo , Sangyeop Lee , Jaehong Cho , Hyunmin Choi , Sanghyeon Lee , Hyungkyu Ham , Gwangsun Kim , Divya Mahajan , Jongse Park View a PDF of the paper titled NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM Inferencing, by Guseul Heo and 8 other authors View PDF Abst
Explore this link on the map →saved by
related reading
- [2312.15159] Understanding the Potential of FPGA-Based Spatial Acceleration for Large Language Model Inferencearxiv.org
- [2403.06664] Smart-Infinity: Fast Large Language Model Training using Near-Storage Processing on a Real Systemarxiv.org
- HeteroLLM: Accelerating Large Language Model Inference on Mobile SoCs with Heterogeneous AI Acceleratorsarxiv.org
- Defeating Nondeterminism in LLM Inference - Thinking Machines Labthinkingmachines.ai
- How To Scale Your Modeljax-ml.github.io
- Transformer Inference Arithmetic | kipply's blogkipp.ly
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- LLM Inference Performance Engineering: Best Practices | Databricks Blogdatabricks.com
- ⭐️ Fast LLM Inference From Scratchandrewkchan.dev
- Inside NVIDIA GPUs: Anatomy of high performance matmul kernels - Aleksa Gordićaleksagordic.com
- Accelerated Inference for Large Transformer Models Using NVIDIA Triton Inference Server | NVIDIA Technical Blogdeveloper.nvidia.com
- PiTorch: ML on Baremetal Raspberry Pis | projectsmasonjwang.com