flâneur — a map of the web's best reading

What Makes Low-Bit Quantization-Aware Training Work for Reasoning LLMs? A Systematic Study

arxiv.org · 8,339 words · saved by 1 readers

This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions. Reasoning models excel at complex tasks such as coding and mathematics, yet their inference is often slow and token-inefficient. To improve the inference efficiency, post-training quantization (PTQ) usually comes with the cost of large accuracy drops, especially for reasoning tasks under low-bit settings. In this study, we present a systematic empirical study of quantization-aware training (QAT) for reasoning models. Our key findings include: (1) Knowledge distillation is a robust objective for reasoning models trained via either supervised fine-tuning or reinforcement learning; (2) PTQ provides a strong initialization for QAT, improving accuracy while reducing training cost; (3) Reinforcement learning remains feasible for quantized models given a viable

What Makes Low-Bit Quantization-Aware Training Work for Reasoning LLMs? A Systematic Study Keyu Lv 1∗ , Manyi Zhang 2 , Xiaobo Xia 3 , Jingchen Ni 1 , Shannan Yan 1 , Xianzhi Yu 2 , Lu Hou 2 , Chun Yuan 1# , Haoli Bai 2# 1 Shenzhen International Graduate School, Tsinghua University 2 Huawei Technologies 3 National University of Singapore lvky24@mails.tsinghua.edu.cn yuanc@sz.tsinghua.edu.cn baihaoli@huawei.com Equal contribution. # Corresponding authors. Abstract Reasoning models excel at complex tasks such as coding and mathematics, yet their inference is often slow and token-inefficient. To

Explore this link on the map →

saved by

related reading