flâneur

Lossy Text Compression | Hackaday.io

hackaday.io · 1,550 words · saved by 1 readers

Not a member? You should Sign up. Already have an account? Log in. To make the experience fit your profile, pick a username and tell us what interests you. We found and based on your interests. Choose more interests. This project was created on 05/08/2015 and last updated 9 years ago. The lossy text compressor consists of a Perl script and accompanying thesaurus. On startup, the script opens the thesaurus file, and constructs a hash of "word" to "shortest synonym of that word". Some words are filtered out, e.g. abbreviations and 2-letter synonyms, because those don't meet quality standards. After this is simple substitution of every word in the supplied text. This should result in a reduction of file size, while still

Lossless compression techniques can only take you so far. Modern lossy compression techniques, like those used in JPEG or MP3, are able to achieve such high ratios because they permanently discard information. Perceptual coding algorithms try to make changes that will go unnoticed by humans. Let's apply similar coding techniques to traditional text files. Where a typical Lempel-Ziv employs a dictionary, we will instead use a thesaurus. The lossy text compressor consists of a Perl script and accompanying thesaurus. On startup, the script opens the thesaurus file, and constructs a hash of…

saved by

related reading