[2312.15159] Understanding the Potential of FPGA-Based Spatial Acceleration for Large Language Model Inference
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs. arXiv Operational Status Get status notifications via email or slack
[2312.15159] Understanding the Potential of FPGA-Based Spatial Acceleration for Large Language Model Inference --> Computer Science > Machine Learning arXiv:2312.15159 (cs) [Submitted on 23 Dec 2023 ( v1 ), last revised 7 Apr 2024 (this version, v2)] Title: Understanding the Potential of FPGA-Based Spatial Acceleration for Large Language Model Inference Authors: Hongzheng Chen , Jiahao Zhang , Yixiao Du , Shaojie Xiang , Zichao Yue , Niansong Zhang , Yaohui Cai , Zhiru Zhang View a PDF of the paper titled Understanding the Potential of FPGA-Based Spatial Acceleration for Large Language Model I
Explore this link on the map →saved by
related reading
- [2403.06664] Smart-Infinity: Fast Large Language Model Training using Near-Storage Processing on a Real Systemarxiv.org
- [2403.00579] NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM Inferencingarxiv.org
- [2209.10797] DFX: A Low-latency Multi-FPGA Appliance for Accelerating Transformer-based Text Generationarxiv.org
- [2304.07493] OliVe: Accelerating Large Language Models via Hardware-friendly Outlier-Victim Pair Quantizationarxiv.org
- How To Scale Your Modeljax-ml.github.io
- LLM Inference Performance Engineering: Best Practices | Databricks Blogdatabricks.com
- ⭐️ Fast LLM Inference From Scratchandrewkchan.dev
- Optimizing inference · Hugging Facehuggingface.co
- Best practices to accelerate inference for large-scale production workloadstogether.ai
- Transformer Inference Arithmetic | kipply's blogkipp.ly
- How is LLaMa.cpp possible?finbarr.ca
- Look Ma, No Bubbles! Designing a Low-Latency Megakernel for Llama-1B · Hazy Researchhazyresearch.stanford.edu