Quantization and Hardware Architecture Co-Design for Matrix-Vector Multiplications of Large Language Models | IEEE Journals & Magazine | IEEE Xplore
2022 4th International Youth Conference on Radio Electronics, Electrical and Power Engineering (REEPE) Published: 2022 About IEEE Xplore | Contact Us | Help | Accessibility | Terms of Use | Nondiscrimination Policy | IEEE Ethics Reporting | Sitemap | IEEE Privacy Policy A not-for-profit organization, IEEE is the world's largest technical professional organization dedicated to advancing technology for the benefit of humanity. © Copyright 2024 IEEE - All rights reserved. A not-for-profit organization, IEEE is the world's largest technical professional organization dedicated to advancing technology for the benefit of humanity.© Copyright 2024 IEEE - All rights reserved. Use of this web site signifies your agreement to the terms and conditions.
Download PDF Download References Request Permissions Save to Alerts Abstract: Large language models (LLMs) have sparked a new revolution in the field of natural language processing (NLP), and have garnered tremendous attention in both academic rese...Show More Metadata Abstract: Large language models (LLMs) have sparked a new revolution in the field of natural language processing (NLP), and have garnered tremendous attention in both academic research and everyday life, thanks to their unprecedented performance in a wide range of applications. However, their deployment remains a…
saved by
related reading
- [2304.07493] OliVe: Accelerating Large Language Models via Hardware-friendly Outlier-Victim Pair Quantizationarxiv.org
- AttentionLego: An Open-Source Building Block For Spatially-Scalable Large Language Model Accelerator With Processing-In-Memory Technologyarxiv.org
- [2312.15159] Understanding the Potential of FPGA-Based Spatial Acceleration for Large Language Model Inferencearxiv.org
- [2403.00579] NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM Inferencingarxiv.org
- Chameleon: a Heterogeneous and Disaggregated Accelerator System for Retrieval-Augmented Language Modelsarxiv.org
- [2403.06664] Smart-Infinity: Fast Large Language Model Training using Near-Storage Processing on a Real Systemarxiv.org
- A Guide to Quantization in LLMs | Symbl.aisymbl.ai
- Quantization from the ground upngrok.com
- Scalable MatMul-free Language Modelingarxiv.org
- How To Scale Your Modeljax-ml.github.io
- AWQ: Activation-aware Weight Quantization for On-Device LLM Compression and Accelerationarxiv.org
- SmoothQuant: Accurate and EfficientPost-Training Quantization for Large Language Modelsarxiv.org