The sample efficiency black hole - by Dwarkesh Patel
"We see these AIs as a galaxy glittering with capabilities, but at their center, invisible to the naked eye, holding all the constellations together, is an unimaginably massive black hole of data."
One definition of intelligence is sample efficiency - that is to say, how much data do you need to see in a given domain in order to operate fluently and competently. It’s not clear that we’ve actually made much progress on training sample efficiency over the last few years - it seems like more so we’ve dramatically widened and improved the data distribution. The main way that AIs have been getting better is from adding more and better data, and scaling the compute to develop that data in the first place. Obviously RL is the main way that has happened. You can think of RL as a kind of…
saved by
related reading
- Sweatshop data is overmechanize.work
- The Future of Meta Superintelligence: A 1 Year Progress Updatenewsletter.semianalysis.com
- Will scaling work?dwarkeshpatel.com
- The Only Important Technology Is The Internet - Kevin Lukevinlu.ai
- Data bottlenecks won’t prevent an intelligence explosionnewsletter.forethought.org
- The Extreme Inefficiency of RL for Frontier Models - Toby Ordtobyord.com
- There Are No New Ideas in AI… Only New Datasetsblog.jxmo.io
- Ilya Sutskever — We're moving from the age of scaling to the age of researchdwarkesh.com
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- Dario Amodei — "We are near the end of the exponential"dwarkesh.com
- Questions about the Future of AI - by Dwarkesh Pateldwarkesh.com
- The real data wall is billions of years of evolutiondynomight.net