✳flâneur — a map of the web's best reading
A Visual Guide to Quantization - by Maarten Grootendorst
newsletter.maartengrootendorst.com · 4,482 words · saved by 1 readers
Exploring memory-efficient techniques for LLMs
A Visual Guide to Quantization Demystifying the Compression of Large Language Models Maarten Grootendorst Jul 22, 2024 524 26 44 Share Translations - Korean - Chinese - French As their name suggests, Large Language Models (LLMs) are often too large to run on consumer hardware. These models may exceed billions of parameters and generally need GPUs with large amounts of VRAM to speed up inference. As such, more and more research has been focused on making these models smaller through improved training, adapters, etc. One major technique in this field is called quantization . In this post, I will
Explore this link on the map →related reading
- A Guide to Quantization in LLMs | Symbl.aisymbl.ai
- What Makes Low-Bit Quantization-Aware Training Work for Reasoning LLMs? A Systematic Studyarxiv.org
- SmoothQuant: Accurate and EfficientPost-Training Quantization for Large Language Modelsarxiv.org
- Integer Quantization for Deep Learning Inference: Principles and Empirical Evaluationarxiv.org
- Quantization · Hugging Facehuggingface.co
- Achieving FP32 Accuracy for INT8 Inference Using Quantization Aware Training with NVIDIA TensorRT | NVIDIA Technical Blogdeveloper.nvidia.com
- LLM.int8()arxiv.org
- [2306.07629] SqueezeLLM: Dense-and-Sparse Quantizationarxiv.org
- LoRA Without Regret - Thinking Machines Labthinkingmachines.ai
- In-depth guide to fine-tuning LLMs with LoRA and QLoRA | Mercity Researchmercity.ai
- [1806.08342] Quantizing deep convolutional networks for efficient inference: A whitepaperarxiv.org
- The 4-bitter Lesson | humans&humansand.ai