Do Qwen3.6 27B quantizations break the pelican? - Quesma Blog
quesma.com · 1,671 words · saved by 1 readers
We tested Qwen3.6 27B quantizations by Unsloth on Hugging Face, with pelicans on bikes, gears, Terminal-Bench 2.1, and AIME-120.
How does model quantization impact its quality? Could you guess which Qwen3.6 27B quantization was used? The largest 8-bit, the smallest 2-bit, or something in between? Not long ago, I was raving about Qwen3.6 27B. The Hacker News discussion was about whether you need a beast or can use a much smaller machine. Even an 8-bit model required less than 45 GB of RAM - a lot for an NVIDIA card (an RTX 5090 has 32 GB), but doable for a strong Apple silicon laptop. But with a smaller quantization we can go much smaller - less than 30 GB for Q4_K_M, and less than 18 GB for the smallest 2-bit…
saved by
related reading
- Quantization from the ground upngrok.com
- Efficient LLM inferencefinbarrtimbers.substack.com
- A Guide to Quantization in LLMs | Symbl.aisymbl.ai
- 2305.14314arxiv.org
- What Makes Low-Bit Quantization-Aware Training Work for Reasoning LLMs? A Systematic Studyarxiv.org
- Composer2.pdfcursor.com
- The 4-bitter Lesson | humans&humansand.ai
- The Model You Audit Is Not the Model You Shiptechpolicy.press
- TurboQuant: Redefining AI efficiency with extreme compressionresearch.google
- 2210.17323arxiv.org
- On neural scaling and the quanta hypothesisericjmichaud.com
- A Visual Guide to Quantization - by Maarten Grootendorstnewsletter.maartengrootendorst.com