[2209.10797] DFX: A Low-latency Multi-FPGA Appliance for Accelerating Transformer-based Text Generation
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs. arXiv Operational Status Get status notifications via email or slack
[2209.10797] DFX: A Low-latency Multi-FPGA Appliance for Accelerating Transformer-based Text Generation --> Electrical Engineering and Systems Science > Systems and Control arXiv:2209.10797 (eess) [Submitted on 22 Sep 2022] Title: DFX: A Low-latency Multi-FPGA Appliance for Accelerating Transformer-based Text Generation Authors: Seongmin Hong , Seungjae Moon , Junsoo Kim , Sungjae Lee , Minsub Kim , Dongsoo Lee , Joo-Young Kim View a PDF of the paper titled DFX: A Low-latency Multi-FPGA Appliance for Accelerating Transformer-based Text Generation, by Seongmin Hong and 6 other authors View PDF
saved by
related reading
- FTRANS: Energy-Efficient Acceleration of Transformers using FPGAarxiv.org
- [2312.15159] Understanding the Potential of FPGA-Based Spatial Acceleration for Large Language Model Inferencearxiv.org
- Hardware Accelerator for Multi-Head Attention and Position-Wise Feed-Forward in the Transformerieeexplore.ieee.org
- AttentionLego: An Open-Source Building Block For Spatially-Scalable Large Language Model Accelerator With Processing-In-Memory Technologyarxiv.org
- [2403.00579] NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM Inferencingarxiv.org
- How To Scale Your Modeljax-ml.github.io
- All About Transformer Inferencejax-ml.github.io
- Accelerated Inference for Large Transformer Models Using NVIDIA Triton Inference Server | NVIDIA Technical Blogdeveloper.nvidia.com
- Transformer Explainer: LLM Transformer Model Visually Explainedpoloclub.github.io
- Transformer Inference Arithmetic | kipply's blogkipp.ly
- Efficiently Scaling Transformer Inferencearxiv.org
- [2207.00032] DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scalearxiv.org