[2304.07493] OliVe: Accelerating Large Language Models via Hardware-friendly Outlier-Victim Pair Quantization
Pi Day is finally here – and so is Giving Day for arXiv! Donate today to directly support arXiv initiatives and help keep open science open. Infinite decimals. Infinite new ideas to discover. Infinite reasons to give. Pi Day is Giving Day: arXiv depends on donations to operate and keep science open for all. Give back to arXiv on 3.14.24! Help | Advanced Search arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs. arXiv Operational Status Get status notifications via email or slack
[2304.07493] OliVe: Accelerating Large Language Models via Hardware-friendly Outlier-Victim Pair Quantization --> Computer Science > Hardware Architecture arXiv:2304.07493 (cs) [Submitted on 15 Apr 2023] Title: OliVe: Accelerating Large Language Models via Hardware-friendly Outlier-Victim Pair Quantization Authors: Cong Guo , Jiaming Tang , Weiming Hu , Jingwen Leng , Chen Zhang , Fan Yang , Yunxin Liu , Minyi Guo , Yuhao Zhu View a PDF of the paper titled OliVe: Accelerating Large Language Models via Hardware-friendly Outlier-Victim Pair Quantization, by Cong Guo and 8 other authors View PDF
saved by
related reading
- Quantization and Hardware Architecture Co-Design for Matrix-Vector Multiplications of Large Language Modelsieeexplore.ieee.org
- AttentionLego: An Open-Source Building Block For Spatially-Scalable Large Language Model Accelerator With Processing-In-Memory Technologyarxiv.org
- [2403.06664] Smart-Infinity: Fast Large Language Model Training using Near-Storage Processing on a Real Systemarxiv.org
- [2201.06618] VAQF: Fully Automatic Software-Hardware Co-Design Framework for Low-Bit Vision Transformerarxiv.org
- [2312.15159] Understanding the Potential of FPGA-Based Spatial Acceleration for Large Language Model Inferencearxiv.org
- How To Scale Your Modeljax-ml.github.io
- Quantization from the ground upngrok.com
- A Guide to Quantization in LLMs | Symbl.aisymbl.ai
- Efficient LLM inferencefinbarrtimbers.substack.com
- Transformers Inference Optimization Toolset | AstraBlogastralord.github.io
- SmoothQuant: Accurate and EfficientPost-Training Quantization for Large Language Modelsarxiv.org
- 2305.14314arxiv.org