flâneur — a map of the web's best reading

How we collected 10,000 hours of neuro-language data in our basement - Conduit

condu.it · 7,642 words · saved by 6 readers

Over the last 6 months, we collected ~10k hours of data across thousands of unique individuals. As far as we know, this is the largest neuro-language dataset in the world.[1] See here, here, here, here, and here (discussion only, no data available) for some of the larger datasets. See recent papers discussing the problem of small datasets here, here, and here. Why did we do this? We train thought-to-text models. That is, we train models to decode semantic content from noninvasive neural data. Here are some entirely zero-shot examples: We'll write about the model in a future post. But before you can train a model that generalizes to new people, you need to get many thousands of hours of data. When we started, the existing datasets were either inapplicable or tiny. Most were in the low hundreds of hours (if that), and most had tens or, at a stretch, hundreds of subjects. So we got thousands of people to come wear headsets in our basement. This post is about how we collected our dataset—

How we collected 10,000 hours of neuro-language data in our basement - Conduit conduit Back to Thoughts View All Posts December 7, 2025 How we collected 10,000 hours of neuro-language data in our basement What we learned about operations and ML from running data collection 20 hours a day Nicholas Aldrich, Lydia Nottingham, Rio Popper, Clem von Stengel Over the last 6 months, we collected ~10k hours of data across thousands of unique individuals. As far as we know, this is the largest neuro-language dataset in the world. [1] See here , here , here , here , and here (discussion only, no data ava

Explore this link on the map →

saved by

related reading