What Makes Low-Bit Quantization-Aware Training Work for Reasoning LLMs? A Systematic Study
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions. Reasoning models excel at complex tasks such as coding and mathematics, yet their inference is often slow and token-inefficient. To improve the inference efficiency, post-training quantization (PTQ) usually comes with the cost of large accuracy drops, especially for reasoning tasks under low-bit settings. In this study, we present a systematic empirical study of quantization-aware training (QAT) for reasoning models. Our key findings include: (1) Knowledge distillation is a robust objective for reasoning models trained via either supervised fine-tuning or reinforcement learning; (2) PTQ provides a strong initialization for QAT, improving accuracy while reducing training cost; (3) Reinforcement learning remains feasible for quantized models given a viable
What Makes Low-Bit Quantization-Aware Training Work for Reasoning LLMs? A Systematic Study Keyu Lv 1∗ , Manyi Zhang 2 , Xiaobo Xia 3 , Jingchen Ni 1 , Shannan Yan 1 , Xianzhi Yu 2 , Lu Hou 2 , Chun Yuan 1# , Haoli Bai 2# 1 Shenzhen International Graduate School, Tsinghua University 2 Huawei Technologies 3 National University of Singapore lvky24@mails.tsinghua.edu.cn yuanc@sz.tsinghua.edu.cn baihaoli@huawei.com Equal contribution. # Corresponding authors. Abstract Reasoning models excel at complex tasks such as coding and mathematics, yet their inference is often slow and token-inefficient. To
Explore this link on the map →saved by
related reading
- You Only Need Minimal RLVR Training: Extrapolating LLMs via Rank-1 Trajectoriesarxiv.org
- LoRA Without Regret - Thinking Machines Labthinkingmachines.ai
- DeepSeek-R1arxiv.org
- A Guide to Quantization in LLMs | Symbl.aisymbl.ai
- The State of Reinforcement Learning for LLM Reasoningmagazine.sebastianraschka.com
- State of RL for reasoning LLMs | A. Weersaweers.de
- Explore | alphaXivalphaxiv.org
- A Visual Guide to Quantization - by Maarten Grootendorstnewsletter.maartengrootendorst.com
- Composer2.pdfcursor.com
- DeepSeek R1's recipe to replicate o1 and the future of reasoning LMsinterconnects.ai
- the-illusion-of-thinking.pdfml-site.cdn-apple.com
- The State of Reinforcement Learning for LLM Reasoningsebastianraschka.com