[2011.14203] EdgeBERT: Sentence-Level Energy Optimizations for Latency-Aware Multi-Task NLP Inference
Pi Day is finally here – and so is Giving Day for arXiv! Donate today to directly support arXiv initiatives and help keep open science open. Infinite decimals. Infinite new ideas to discover. Infinite reasons to give. Pi Day is Giving Day: arXiv depends on donations to operate and keep science open for all. Give back to arXiv on 3.14.24! Help | Advanced Search arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs. arXiv Operational Status Get status notifications via email or slack
[2011.14203] EdgeBERT: Sentence-Level Energy Optimizations for Latency-Aware Multi-Task NLP Inference --> Computer Science > Hardware Architecture arXiv:2011.14203 (cs) [Submitted on 28 Nov 2020 ( v1 ), last revised 6 Sep 2021 (this version, v5)] Title: EdgeBERT: Sentence-Level Energy Optimizations for Latency-Aware Multi-Task NLP Inference Authors: Thierry Tambe , Coleman Hooper , Lillian Pentecost , Tianyu Jia , En-Yu Yang , Marco Donato , Victor Sanh , Paul N. Whatmough , Alexander M. Rush , David Brooks , Gu-Yeon Wei View a PDF of the paper titled EdgeBERT: Sentence-Level Energy Optimizati
Explore this link on the map →saved by
related reading
- [2312.15159] Understanding the Potential of FPGA-Based Spatial Acceleration for Large Language Model Inferencearxiv.org
- [2012.09852] SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head Pruningarxiv.org
- [2209.10797] DFX: A Low-latency Multi-FPGA Appliance for Accelerating Transformer-based Text Generationarxiv.org
- [2201.06618] VAQF: Fully Automatic Software-Hardware Co-Design Framework for Low-Bit Vision Transformerarxiv.org
- How To Scale Your Modeljax-ml.github.io
- Transformer Inference Arithmetic | kipply's blogkipp.ly
- [2101.01321] I-BERT: Integer-only BERT Quantizationarxiv.org
- [2005.03842] GOBO: Quantizing Attention-Based NLP Models for Low Latency and Energy Efficient Inferencearxiv.org
- Transformers Inference Optimization Toolset | AstraBlogastralord.github.io
- LLM Inference Performance Engineering: Best Practices | Databricks Blogdatabricks.com
- Overleaf Examplearxiv.org
- Best practices to accelerate inference for large-scale production workloadstogether.ai