The Model You Audit Is Not the Model You Ship | TechPolicy.Press
techpolicy.press · 1,105 words · saved by 1 readers
Almost every LLM that reaches a broad audience is quantized—with under-appreciated ramifications for AI safety, writes Emilio Ferrara.
Linus Zoll & Google DeepMind / Better Images of AI / Generative Image models / CC-BY 4.0 Almost every large language model that reaches a broad audience is quantized. It is trained at full precision, then compressed so it can run affordably on a phone, on a laptop, on a server rack a fraction of the size originally needed. Compression is what makes local and low-cost AI economically possible, and it is spreading quickly. It is also treated as a nonevent for safety. The model is evaluated at full precision, the compressed build inherits the evaluation, and only rarely does anyone recheck.…
saved by
related reading
- A Safe Path to Open Weights - Thinking Machines Labthinkingmachines.ai
- Quantization from the ground upngrok.com
- AI in 2025: gestalt — LessWronglesswrong.com
- A Guide to Quantization in LLMs | Symbl.aisymbl.ai
- Efficient LLM inferencefinbarrtimbers.substack.com
- 2305.14314arxiv.org
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- Do Qwen3.6 27B quantizations break the pelican?quesma.com
- [2506.17209] Fine-Tuning Lowers Safety and Disrupts Evaluation Consistencyarxiv.org
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Fine-Tuning Lowers Safety and Disrupts Evaluation Consistencyarxiv.org
- What Makes Low-Bit Quantization-Aware Training Work for Reasoning LLMs? A Systematic Studyarxiv.org