flâneur

Do Qwen3.6 27B quantizations break the pelican? - Quesma Blog

quesma.com · 1,671 words · saved by 1 readers

We tested Qwen3.6 27B quantizations by Unsloth on Hugging Face, with pelicans on bikes, gears, Terminal-Bench 2.1, and AIME-120.

How does model quantization impact its quality? Could you guess which Qwen3.6 27B quantization was used? The largest 8-bit, the smallest 2-bit, or something in between? Not long ago, I was raving about Qwen3.6 27B. The Hacker News discussion was about whether you need a beast or can use a much smaller machine. Even an 8-bit model required less than 45 GB of RAM - a lot for an NVIDIA card (an RTX 5090 has 32 GB), but doable for a strong Apple silicon laptop. But with a smaller quantization we can go much smaller - less than 30 GB for Q4_K_M, and less than 18 GB for the smallest 2-bit…

saved by

related reading