[2201.06618] VAQF: Fully Automatic Software-Hardware Co-Design Framework for Low-Bit Vision Transformer
Pi Day is finally here – and so is Giving Day for arXiv! Donate today to directly support arXiv initiatives and help keep open science open. Infinite decimals. Infinite new ideas to discover. Infinite reasons to give. Pi Day is Giving Day: arXiv depends on donations to operate and keep science open for all. Give back to arXiv on 3.14.24! Help | Advanced Search arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs. arXiv Operational Status Get status notifications via email or slack
[2201.06618] VAQF: Fully Automatic Software-Hardware Co-Design Framework for Low-Bit Vision Transformer --> Computer Science > Machine Learning arXiv:2201.06618 (cs) [Submitted on 17 Jan 2022 ( v1 ), last revised 18 Feb 2022 (this version, v2)] Title: VAQF: Fully Automatic Software-Hardware Co-Design Framework for Low-Bit Vision Transformer Authors: Mengshu Sun , Haoyu Ma , Guoliang Kang , Yifan Jiang , Tianlong Chen , Xiaolong Ma , Zhangyang Wang , Yanzhi Wang View a PDF of the paper titled VAQF: Fully Automatic Software-Hardware Co-Design Framework for Low-Bit Vision Transformer, by Mengshu
Explore this link on the map →saved by
related reading
- [2304.07493] OliVe: Accelerating Large Language Models via Hardware-friendly Outlier-Victim Pair Quantizationarxiv.org
- [2011.14203] EdgeBERT: Sentence-Level Energy Optimizations for Latency-Aware Multi-Task NLP Inferencearxiv.org
- [2012.09852] SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head Pruningarxiv.org
- [2312.15159] Understanding the Potential of FPGA-Based Spatial Acceleration for Large Language Model Inferencearxiv.org
- [2209.10797] DFX: A Low-latency Multi-FPGA Appliance for Accelerating Transformer-based Text Generationarxiv.org
- [2401.02721] A Cost-Efficient FPGA Implementation of Tiny Transformer Model using Neural ODEarxiv.org
- A friendly introduction to machine learning compilers and optimizershuyenchip.com
- Transformers Inference Optimization Toolset | AstraBlogastralord.github.io
- [2302.05442] Scaling Vision Transformers to 22 Billion Parametersarxiv.org
- Transformer Math 101 | EleutherAI Blogblog.eleuther.ai
- Accelerated Inference for Large Transformer Models Using NVIDIA Triton Inference Server | NVIDIA Technical Blogdeveloper.nvidia.com
- [2605.05331] ViTok-v2: Scaling Native Resolution Auto-Encoders to 5 Billion Parametersarxiv.org