AWQ: Activation-aware Weight Quantization for On-Device LLM Compression and Acceleration
arxiv.org · 6,251 words · saved by 1 readers
N/A
AWQ: A CTIVATION - AWARE W EIGHT Q UANTIZATION FOR O N -D EVICE LLM C OMPRESSION AND ACCELERATION Ji Lin * 1 Jiaming Tang * 1 2 Haotian Tang † 1 Shang Yang † 1 Wei-Ming Chen 3 Wei-Chen Wang 1 Guangxuan Xiao 1 Xingyu Dang 1 4 Chuang Gan 5 6 Song Han 1 3 https://github.com/mit-han-lab/llm-awq…
related reading
- A Guide to Quantization in LLMs | Symbl.aisymbl.ai
- Quantization from the ground upngrok.com
- Efficient LLM inferencefinbarrtimbers.substack.com
- Quantization and Hardware Architecture Co-Design for Matrix-Vector Multiplications of Large Language Modelsieeexplore.ieee.org
- SmoothQuant: Accurate and EfficientPost-Training Quantization for Large Language Modelsarxiv.org
- LLM Inference Performance Engineering: Best Practices | Databricks Blogdatabricks.com
- A Visual Guide to Quantization - by Maarten Grootendorstnewsletter.maartengrootendorst.com
- 2408.14690arxiv.org
- The Big LLM Architecture Comparisonmagazine.sebastianraschka.com
- 2305.14314arxiv.org
- How is LLaMa.cpp possible?finbarr.ca
- What Makes Low-Bit Quantization-Aware Training Work for Reasoning LLMs? A Systematic Studyarxiv.org