[2401.11851] BETA: Binarized Energy-Efficient Transformer Accelerator at the Edge
Pi Day is finally here – and so is Giving Day for arXiv! Donate today to directly support arXiv initiatives and help keep open science open. Infinite decimals. Infinite new ideas to discover. Infinite reasons to give. Pi Day is Giving Day: arXiv depends on donations to operate and keep science open for all. Give back to arXiv on 3.14.24! Help | Advanced Search arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs. arXiv Operational Status Get status notifications via email or slack
View PDF HTML (experimental) Abstract:Existing binary Transformers are promising in edge deployment due to their compact model size, low computational complexity, and considerable inference accuracy. However, deploying binary Transformers faces challenges on prior processors due to inefficient execution of quantized matrix multiplication (QMM) and the energy consumption overhead caused by multi-precision activations. To tackle the challenges above, we first develop a computation flow abstraction method for binary Transformers to improve QMM execution efficiency by optimizing the computation…
saved by
related reading
- FTRANS: Energy-Efficient Acceleration of Transformers using FPGAarxiv.org
- [2201.06618] VAQF: Fully Automatic Software-Hardware Co-Design Framework for Low-Bit Vision Transformerarxiv.org
- [2011.14203] EdgeBERT: Sentence-Level Energy Optimizations for Latency-Aware Multi-Task NLP Inferencearxiv.org
- AttentionLego: An Open-Source Building Block For Spatially-Scalable Large Language Model Accelerator With Processing-In-Memory Technologyarxiv.org
- [2403.00579] NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM Inferencingarxiv.org
- Hardware Accelerator for Multi-Head Attention and Position-Wise Feed-Forward in the Transformerieeexplore.ieee.org
- [2304.07493] OliVe: Accelerating Large Language Models via Hardware-friendly Outlier-Victim Pair Quantizationarxiv.org
- Transformer Inference Arithmetic | kipply's blogkipp.ly
- [2401.02721] A Cost-Efficient FPGA Implementation of Tiny Transformer Model using Neural ODEarxiv.org
- The Annotated Transformernlp.seas.harvard.edu
- PiTorch: ML on Baremetal Raspberry Pis | projectsmasonjwang.com
- Transformer Math 101 | EleutherAI Blogblog.eleuther.ai