Silent speech with ultrasound — Aleph
We trained a model to predict speech from ultrasound video of the tongue, no speaking required. It reads open text, generalizes across people, and reaches a 15.6% word error rate on open-vocabulary speech.
Silent speech with ultrasound Vadims Casecnikovs ✉ , Gimran Abdullin , Raffi Hotter, Lev Chizhov July 7, 2026 For comparison, lip-reading achieves 12.5% word error rate on a 1M hour dataset. We're excited that our system approaches existing methods despite being an early investigation trained on a 50-hour dataset and done in just a month. We place an ultrasound probe behind the chin and capture videos like this of the tongue: Your browser does not support the video tag. And then we turn them into words. Here's a quick demo: Your browser does not support the video tag. A few days ago, we found
Explore this link on the map →saved by
related reading
- Crossing the uncanny valley of conversational voice | Sesamesesame.com
- Introducing talkie: a 13B vintage language model from 1930talkie-lm.com
- [2105.11084] Unsupervised Speech Recognitionarxiv.org
- Reproducing DeepTFUS | projectsmasonjwang.com
- How we collected 10,000 hours of neuro-language data in our basement - Conduitcondu.it
- Tincans - Final Technical Reporttincans.ai
- whisper/model-card.md at main · openai/whisper · GitHubgithub.com
- Unreal Speech: Cheapest Text-to-Speech APIunrealspeech.com
- Inkling: Our Open-Weights Model - Thinking Machines Labthinkingmachines.ai
- Subvocal recognition - Wikipediaen.wikipedia.org
- GitHub - openai/whisper: Robust Speech Recognition via Large-Scale Weak Supervision · GitHubgithub.com
- Verifying your browser | OpenReviewopenreview.net