flâneur — a map of the web's best reading

marcodsn.me/posts/exploring-mimi

marcodsn.me · saved by 1 readers

Today we are going to explore Mimi, a state-of-the-art neural audio codec developed by Kyutai. We will compare audio quality before and after Mimi processing, and we will then study how Mimi works and what types of outputs it produces. Mimi codec is a state-of-the-art audio neural codec, developed by Kyutai, that combines semantic and acoustic information into audio tokens running at 12Hz and a bitrate of 1.1kbps. This is the official description of Mimi from the Kyutai’s Huggingface repo. Mimi is also the neural audio codec that powers Moshi (check moshi.chat for a demo), which we may discuss in a future post. First, let’s define a codec: A codec, short for “coder-decoder” or “compressor-decompressor”, is a device or computer program that encodes or decodes a digital data stream or signal. The primary purposes of codecs are to: An audio codec is a codec that encodes or decodes audio data. There exist two types of audio codecs: A neural audio codec is a type of lossy audio codec that l

Today we are going to explore Mimi, a state-of-the-art neural audio codec developed by Kyutai. We will compare audio quality before and after Mimi processing, and we will then study how Mimi works and what types of outputs it produces. Mimi codec is a state-of-the-art audio neural codec, developed by Kyutai, that combines semantic and acoustic information into audio tokens running at 12Hz and a bitrate of 1.1kbps. This is the official description of Mimi from the Kyutai’s Huggingface repo. Mimi is also the neural audio codec that powers Moshi (check moshi.chat for a demo), which we may discuss

Explore this link on the map →