[2209.10797] DFX: A Low-latency Multi-FPGA Appliance for Accelerating Transformer-based Text Generation
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs. arXiv Operational Status Get status notifications via email or slack
[2209.10797] DFX: A Low-latency Multi-FPGA Appliance for Accelerating Transformer-based Text Generation --> Electrical Engineering and Systems Science > Systems and Control arXiv:2209.10797 (eess) [Submitted on 22 Sep 2022] Title: DFX: A Low-latency Multi-FPGA Appliance for Accelerating Transformer-based Text Generation Authors: Seongmin Hong , Seungjae Moon , Junsoo Kim , Sungjae Lee , Minsub Kim , Dongsoo Lee , Joo-Young Kim View a PDF of the paper titled DFX: A Low-latency Multi-FPGA Appliance for Accelerating Transformer-based Text Generation, by Seongmin Hong and 6 other authors View PDF
Explore this link on the map →saved by
related reading
- [2312.15159] Understanding the Potential of FPGA-Based Spatial Acceleration for Large Language Model Inferencearxiv.org
- [2403.00579] NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM Inferencingarxiv.org
- [2011.14203] EdgeBERT: Sentence-Level Energy Optimizations for Latency-Aware Multi-Task NLP Inferencearxiv.org
- How To Scale Your Modeljax-ml.github.io
- Transformer Inference Arithmetic | kipply's blogkipp.ly
- Transformer Explainer: LLM Transformer Model Visually Explainedpoloclub.github.io
- Accelerated Inference for Large Transformer Models Using NVIDIA Triton Inference Server | NVIDIA Technical Blogdeveloper.nvidia.com
- [2207.00032] DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scalearxiv.org
- Assisted Generation: a new direction toward low-latency text generationhuggingface.co
- Transformers Inference Optimization Toolset | AstraBlogastralord.github.io
- Accelerating Generative AI with PyTorch II: GPT, Fast – PyTorchpytorch.org
- The Annotated Transformernlp.seas.harvard.edu