AI and the 'Why Now' of Data DAOs – Variant
Recent high-profile data licensing deals such as those between OpenAI and News Corp and Reddit underscore the need for high-quality data in AI. Frontier models are already trained on much of the internet—for example, Common Crawl, which indexes about 10% of all web pages, is used for LLM training and contains over 100 trillion tokens. An avenue for further improvement in AI models is to expand and enhance the data they can train on. We’ve been discussing mechanisms for how data could be aggregated—particularly in a decentralized way. We’re especially interested in exploring how decentralized methods could help generate new datasets and economically reward contributors and creators. One topic of discussion within crypto in the last few years is the idea of data DAOs, or collectives of individuals who create, organize, and govern data. The topic has been covered by Multicoin and others, but the rapid advancement of AI is a catalyst for a new “why now?” of data DAOs. We wanted to share ou
AI and the 'Why Now' of Data DAOs – Variant AI and the ‘Why Now’ of Data DAOs Data DAOs represent one path to generating new high-quality data sets and overcoming the data wall in AI June 18, 2024 by Li Jin --> This post first appeared on Li’s blog . Recent high-profile data licensing deals such as those between OpenAI and News Corp and Reddit underscore the need for high-quality data in AI. Frontier models are already trained on much of the internet—for example, Common Crawl, which indexes about 10% of all web pages, is used for LLM training and contains over
related reading
- The Only Important Technology Is The Internet - Kevin Lukevinlu.ai
- A Stargate for Data - by willdepue - Will DePuewilldepue.substack.com
- There Are No New Ideas in AI… Only New Datasetsblog.jxmo.io
- Unlocking a Million Times More Data for AI | IFPifp.org
- Is Synthetic Data the Key to AGI?digitalspirits.substack.com
- State of Data (Jan 2026)seancai.com
- Sweatshop data is overmechanize.work
- will depue on X: "A Stargate for Data Labs are on a trajectory towards >$100B/year of data spend by 2030. As we begin the trillion-dollar compute project, we need to think about the equivalent civilizational-scale effort for the other core x.com
- A World of Verifiable Domainsseancai.com
- Training Imperativesdan.io
- Gaming Worlds Could Be The Answer To AI’s Data Problemforbes.com
- Defensibility in the Age of AIsethbannon.com