Background — Connecting Music Audio and Natural Language
Fig. 2 Image Source: Learning the Meaning of Music by Brian A. Whitman, 2005, Massachusetts Institute of Technology(MIT) The journey of Music and Language Models started with two basic human desires: to understand music deeply and to listen to the music we want whenever we want, whether it’s existing artist music or creative new music. These fundamental needs have driven the development of technologies that connect music and language. This is because language is the most fundamental communication channel we use, and through this language, we aim to communicate with machines. The first approach was Supervised Classification. This method involved developing models that could predict appropriate Natural Language Labels (Fixed-Vocabulary) for given audio inputs. These labels could cover a wide range of musical attributes including genre, mood, style, instruments, usage, theme, key, tempo, and more [SLC07]. The advantage of Supervised Classification was that it automated the annotation proc
Background # Fig. 2 Image Source: Learning the Meaning of Music by Brian A. Whitman, 2005, Massachusetts Institute of Technology(MIT) # The journey of Music and Language Models started with two basic human desires: to understand music deeply and to listen to the music we want whenever we want, whether it’s existing artist music or creative new music. These fundamental needs have driven the development of technologies that connect music and language. This is because language is the most fundamental communication channel we use, and through this language, we aim to communicate with machines. Ear
saved by
related reading
- Generating music in the waveform domain – Sander Dielemansander.ai
- Generating audio for videodeepmind.google
- What are Diffusion Models? | Lil'Loglilianweng.github.io
- MusicLM: Generating Music From Textgoogle-research.github.io
- NotaGenelectricalexis.github.io
- Crossing the uncanny valley of conversational voice | Sesamesesame.com
- The AI Sound Housemaximevidal.com
- radford2018improving.pdfcs.ubc.ca
- What are Diffusion Models?lilianweng.github.io
- Yang Songyang-song.net
- Woosh: A Sound Effects Foundation Modelarxiv.org
- arxiv.org/pdf/1704.01444arxiv.org