flâneur

Quantization from the ground up | ngrok blog

ngrok.com · 6,054 words · saved by 3 readers

A complete guide to what quantization is, how it works, and how it's used to compress large language models

Qwen-3-Coder-Next is an 80 billion parameter model 159.4GB in size. That’s roughly how much RAM you would need to run it, and that’s before thinking about long context windows. This is not considered a big model. Rumors have it that frontier models have over 1 trillion parameters, which would require at least 2TB of RAM. The last time I saw that much RAM in one machine was never. But what if I told you we can make LLMs 4x smaller and 2x faster, enough to run very capable models on your laptop, all while losing only 5-10% accuracy. That’s the magic of quantization. What makes large…

saved by

related reading