Lossy Text Compression | Hackaday.io
Not a member? You should Sign up. Already have an account? Log in. To make the experience fit your profile, pick a username and tell us what interests you. We found and based on your interests. Choose more interests. This project was created on 05/08/2015 and last updated 9 years ago. The lossy text compressor consists of a Perl script and accompanying thesaurus. On startup, the script opens the thesaurus file, and constructs a hash of "word" to "shortest synonym of that word". Some words are filtered out, e.g. abbreviations and 2-letter synonyms, because those don't meet quality standards. After this is simple substitution of every word in the supplied text. This should result in a reduction of file size, while still
Lossless compression techniques can only take you so far. Modern lossy compression techniques, like those used in JPEG or MP3, are able to achieve such high ratios because they permanently discard information. Perceptual coding algorithms try to make changes that will go unnoticed by humans. Let's apply similar coding techniques to traditional text files. Where a typical Lempel-Ziv employs a dictionary, we will instead use a thesaurus. The lossy text compressor consists of a Perl script and accompanying thesaurus. On startup, the script opens the thesaurus file, and constructs a hash of…
saved by
related reading
- ChatGPT Is a Blurry JPEG of the Web | The New Yorkernewyorker.com
- Lossless compressionen.wikipedia.org
- Data Compression Explainedmattmahoney.net
- [2309.10668] Language Modeling Is Compressionarxiv.org
- Compression and Intelligencegreene.sh
- Fabrice Bellard's Home Pagebellard.org
- Compression is predictionngrok.com
- LLMs can invent their own compression - Rajan Agarwalrajan.sh
- How does image compression work? | Lee Robinsonleerob.com
- How Lossless Data Compression Works | Quanta Magazinequantamagazine.org
- Can gzip be a language model?nathan.rs
- 500'000€ Prize for Compressing Human Knowledgeprize.hutter1.net