[2105.11084] Unsupervised Speech Recognition
Despite rapid progress in the recent past, current speech recognition systems still require labeled training data which limits this technology to a small fraction of the languages spoken around the globe. This paper describes wav2vec-U, short for wav2vec Unsupervised, a method to train speech recognition models without any labeled data. We leverage self-supervised speech representations to segment unlabeled audio and learn a mapping from these representations to phonemes via adversarial training. The right representations are key to the success of our method. Compared to the best previous unsupervised work, wav2vec-U reduces the phoneme error rate on the TIMIT benchmark from 26.1 to 11.3. On the larger English Librispeech benchmark, wav2vec-U achieves a word error rate of 5.9 on test-other, rivaling some of the best published systems trained on 960 hours of labeled data from only two years ago. We also experiment on nine other languages, including low-resource languages such as Kyrgyz, Swahili and Tatar.
Unsupervised Speech Recognition Alexei Baevski4 , Wei-Ning Hsu4 , Alexis Conneau∗, Michael Auli4 4 Facebook AI Google AI Abstract Despite rapid progress in the recent past, current speech recognition systems still require arXiv:2105.11084v3 [cs.CL] 2 May 2022 labeled training data which limits this…
related reading
- Silent speech with ultrasound — Alephalephneuro.com
- Crossing the uncanny valley of conversational voice | Sesamesesame.com
- SpecAugment: A New Data Augmentation Method for Automatic Speech Recognitionai.googleblog.com
- Unsupervised Elicitation of Language Modelsarxiv.org
- Learning with not Enough Data Part 1: Semi-Supervised Learning | Lil'Loglilianweng.github.io
- 1301.3781arxiv.org
- radford2018improving.pdfcs.ubc.ca
- arxiv.org/pdf/1704.01444arxiv.org
- WaveNetdeepmind.com
- Unsupervised Elicitationalignment.anthropic.com
- Audio Deep Learning Made Simple: Automatic Speech Recognition (ASR), How it Works | Towards Data Sciencetowardsdatascience.com
- whisper/model-card.md at main · openai/whisper · GitHubgithub.com