flâneur

The Model You Audit Is Not the Model You Ship | TechPolicy.Press

techpolicy.press · 1,105 words · saved by 1 readers

Almost every LLM that reaches a broad audience is quantized—with under-appreciated ramifications for AI safety, writes Emilio Ferrara.

Linus Zoll & Google DeepMind / Better Images of AI / Generative Image models / CC-BY 4.0 Almost every large language model that reaches a broad audience is quantized. It is trained at full precision, then compressed so it can run affordably on a phone, on a laptop, on a server rack a fraction of the size originally needed. Compression is what makes local and low-cost AI economically possible, and it is spreading quickly. It is also treated as a nonevent for safety. The model is evaluated at full precision, the compressed build inherits the evaluation, and only rarely does anyone recheck.…

saved by

related reading