Background — Connecting Music Audio and Natural Language
Fig. 2 Image Source: Learning the Meaning of Music by Brian A. Whitman, 2005, Massachusetts Institute of Technology(MIT) The journey of Music and Language Models started with two basic human desires: to understand music deeply and to listen to the music we want whenever we want, whether it’s existing artist music or creative new music. These fundamental needs have driven the development of technologies that connect music and language. This is because language is the most fundamental communication channel we use, and through this language, we aim to communicate with machines. The first approach was Supervised Classification. This method involved developing models that could predict appropriate Natural Language Labels (Fixed-Vocabulary) for given audio inputs. These labels could cover a wide range of musical attributes including genre, mood, style, instruments, usage, theme, key, tempo, and more [SLC07]. The advantage of Supervised Classification was that it automated the annotation proc
Background # Fig. 2 Image Source: Learning the Meaning of Music by Brian A. Whitman, 2005, Massachusetts Institute of Technology(MIT) # The journey of Music and Language Models started with two basic human desires: to understand music deeply and to listen to the music we want whenever we want, whether it’s existing artist music or creative new music. These fundamental needs have driven the development of technologies that connect music and language. This is because language is the most fundamental communication channel we use, and through this language, we aim to communicate with machines. Ear
Explore this link on the map →saved by
related reading
- Generating music in the waveform domain – Sander Dielemansander.ai
- What are Diffusion Models? | Lil'Loglilianweng.github.io
- NotaGenelectricalexis.github.io
- Crossing the uncanny valley of conversational voice | Sesamesesame.com
- The AI Sound Housemaximevidal.com
- Yang Songyang-song.net
- Woosh: A Sound Effects Foundation Modelarxiv.org
- Language Modelinglena-voita.github.io
- How I Built a Lo-fi Music Web Player with AI-Generated Tracks | Towards Data Sciencetowardsdatascience.com
- To Understand Language is to Understand Generalization | Eric Jangevjang.com
- Generative modelling in latent space – Sander Dielemansander.ai
- Large Language Diffusion Modelsarxiv.org