flâneur — a map of the web's best reading

Silent speech with ultrasound — Aleph

alephneuro.com · 1,293 words · saved by 5 readers

We trained a model to predict speech from ultrasound video of the tongue, no speaking required. It reads open text, generalizes across people, and reaches a 15.6% word error rate on open-vocabulary speech.

Silent speech with ultrasound Vadims Casecnikovs ✉ , Gimran Abdullin , Raffi Hotter, Lev Chizhov July 7, 2026 For comparison, lip-reading achieves 12.5% word error rate on a 1M hour dataset. We're excited that our system approaches existing methods despite being an early investigation trained on a 50-hour dataset and done in just a month. We place an ultrasound probe behind the chin and capture videos like this of the tongue: Your browser does not support the video tag. And then we turn them into words. Here's a quick demo: Your browser does not support the video tag. A few days ago, we found

Explore this link on the map →

saved by

related reading