Is Synthetic Data the Key to AGI? - by Nabeel S. Qureshi
One key fact about modern large language models (LLMs) can be summarized as follows: it’s the dataset, stupid. AI model behavior is largely determined by the dataset it’s trained on; other details, such as architecture, are simply a means of delivering computing power to that dataset. Having a clean, high-quality dataset is worth a lot. [1] The centrality of data is reflected in AI business practice. OpenAI recently announced deals with Axel Springer, Elsevier, the Associated Press, and other publishers and mass media companies for their data; the New York Times (NYT) recently sued OpenAI, demanding that its GPTs, trained on NYT data, be shut down. And Apple is offering $50 million plus for data contracts with publishers. At current margins, models benefit more from additional data than they do from additional size. The expansion in size of training corpora has been rapid. The first modern LLM was trained on Wikipedia. GPT-3 was trained on 300 billion tokens (typically words, parts of
This is part 2 in a series on AI. The first part can be found here. LLMs are trained on vast amounts of data, many libraries’ worth. But what if we run out? Image source: Twitter One key fact about modern large language models (LLMs) can be summarized as follows: it’s the dataset, stupid. AI model behavior is largely determined by the dataset it’s trained on; other details, such as architecture, are simply a means of delivering computing power to that dataset. Having a clean, high-quality dataset is worth a lot. [1] The centrality of data is reflected in AI business practice. OpenAI…
related reading
- There Are No New Ideas in AI… Only New Datasetsblog.jxmo.io
- Will scaling work?dwarkeshpatel.com
- The “it” in AI models is the dataset. — Non_Intnonint.com
- Sporks of AGIsergeylevine.substack.com
- Autodata: An agentic data scientist to create high quality synthetic dataalphaxiv.org
- A Stargate for Data - by willdepue - Will DePuewilldepue.substack.com
- [2409.19759] Balancing Cost and Effectiveness of Synthetic Data Generation Strategies for LLMsarxiv.org
- Things we learned about LLMs in 2024simonwillison.net
- Scaling: The State of Play in AIoneusefulthing.org
- Will scaling work? - by Dwarkesh Patel - Dwarkesh Podcastdwarkesh.com
- chinchilla's wild implications — LessWronglesswrong.com
- Stuff we figured out about AI in 2023simonwillison.net