How we collected 10,000 hours of neuro-language data in our basement - Conduit
Over the last 6 months, we collected ~10k hours of data across thousands of unique individuals. As far as we know, this is the largest neuro-language dataset in the world.[1] See here, here, here, here, and here (discussion only, no data available) for some of the larger datasets. See recent papers discussing the problem of small datasets here, here, and here. Why did we do this? We train thought-to-text models. That is, we train models to decode semantic content from noninvasive neural data. Here are some entirely zero-shot examples: We'll write about the model in a future post. But before you can train a model that generalizes to new people, you need to get many thousands of hours of data. When we started, the existing datasets were either inapplicable or tiny. Most were in the low hundreds of hours (if that), and most had tens or, at a stretch, hundreds of subjects. So we got thousands of people to come wear headsets in our basement. This post is about how we collected our dataset—
How we collected 10,000 hours of neuro-language data in our basement - Conduit conduit Back to Thoughts View All Posts December 7, 2025 How we collected 10,000 hours of neuro-language data in our basement What we learned about operations and ML from running data collection 20 hours a day Nicholas Aldrich, Lydia Nottingham, Rio Popper, Clem von Stengel Over the last 6 months, we collected ~10k hours of data across thousands of unique individuals. As far as we know, this is the largest neuro-language dataset in the world. [1] See here , here , here , here , and here (discussion only, no data ava
Explore this link on the map →saved by
related reading
- Silent speech with ultrasound — Alephalephneuro.com
- POYO-1poyo-brain.github.io
- [2605.15220] Always Learning, Always Mixing: Efficient and Simple Data Mixing All The Timearxiv.org
- On neural scaling and the quanta hypothesisericjmichaud.com
- Crossing the uncanny valley of conversational voice | Sesamesesame.com
- Generalist - GEN-0 / Embodied Foundation Models That Scale with Physical Interactiongeneralistai.com
- Language models can explain neurons in language modelsopenaipublic.blob.core.windows.net
- Fermi estimate of future training runsdanieldewey.net
- chinchilla's wild implications — LessWronglesswrong.com
- 2405.01470arxiv.org
- Why you need to improve your training data, and how to do it << Pete Warden's blogpetewarden.com
- Stanford CS336 | Language Modeling from Scratchcs336.stanford.edu