[2309.10668] Language Modeling Is Compression
Abstract:It has long been established that predictive models can be transformed into lossless compressors and vice versa. Incidentally, in recent years, the machine learning community has focused on training increasingly large and powerful self-supervised (language) models. Since these large language models exhibit impressive predictive capabilities, they are well-positioned to be strong compressors. In this work, we advocate for viewing the prediction problem through the lens of compression and evaluate the compression capabilities of large (foundation) models. We show that large language models are powerful general-purpose predictors and that the compression viewpoint provides novel insights into scaling laws, tokenization, and in-context learning. For example, Chinchilla 70B, while trained primarily on text, compresses ImageNet patches to 43.4% and LibriSpeech samples to 16.4% of their raw size, beating domain-specific compressors like PNG (58.5%) or FLAC (30.3%), respectively. Finally, we show that the prediction-compression equivalence allows us to use any compressor (like gzip) to build a conditional generative model.
# link_129564vty3s.pdf ## Metadata - PDFFormatVersion=1.5 - IsLinearized=false - IsAcroFormPresent=false - IsXFAPresent=false - IsCollectionPresent=false - IsSignaturesPresent=false - CreationDate=D:20240320003032Z - Creator=LaTeX with hyperref - ModDate=D:20240320003032Z - Custom.PTEX.Fullbanner=This is pdfTeX, Version 3.141592653-2.6-1.40.25 (TeX Live 2023) kpathsea version 6.3.5 - Producer=pdfTeX-1.40.25 - Trapped=False ## Contents ### Page 1 Published as a conference paper at ICLR 2024LANGUAGE MODELING IS COMPRESSIONGrégoire Delétang∗1 Anian Ruoss∗1 Paul-Ambroise Duquenne2 Elliot Catt1T
saved by
related reading
- Compression and Intelligencegreene.sh
- Can gzip be a language model?nathan.rs
- Data Compression Explainedmattmahoney.net
- Compression is predictionngrok.com
- LLMs can invent their own compression - Rajan Agarwalrajan.sh
- ChatGPT Is a Blurry JPEG of the Web | The New Yorkernewyorker.com
- On the Tradeoffs of SSMs and Transformers | Goomba Labgoombalab.github.io
- LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Modelsarxiv.org
- [2308.07633] A Survey on Model Compression for Large Language Modelsarxiv.org
- The Intricate Link Between Compression and Predictionmindfulmodeler.substack.com
- A History of Large Language Modelsgregorygundersen.com
- [2005.14165] Language Models are Few-Shot Learnersarxiv.org