The Bitter Lesson is Misunderstood - by Kushal Chakrabarti
tl;dr: For years, we've been reading the Bitter Lesson backwards. It wasn't about compute — it was about data. Here's the part of Scaling Laws no one talks about: Translation: Double your GPUs? You need 40% more data or you're just lighting cash on fire. But there's no 2nd Internet (we’ve already eaten the first one). The path forward: data alchemists (high-variance, 300% lottery ticket) or model architects (20-30% steady gains), not chip buyers. Full analysis below. For almost a decade, the most important essay in AI has been Rich Sutton’s Bitter Lesson. And for years, I think we’ve profoundly misunderstood it. Talking to dozens of researchers at top labs over the past year, I think this might be the single-most common (and single-most dangerous) fallacy in AI today. The lesson, as we all learned it, was a beautifully reductionist gospel: stop trying to be clever. Here’s Sutton’s opening sentence: The biggest lesson that can be read from 70 years of AI research is that general methods
tl;dr: For years, we've been reading the Bitter Lesson backwards. It wasn't about compute — it was about data. Here's the part of Scaling Laws no one talks about: \(C \sim D^2\) Translation: Double your GPUs? You need 40% more data or you're just lighting cash on fire. But there's no 2nd Internet (we’ve already eaten the first one). The path forward: data alchemists (high-variance, 300% lottery ticket) or model architects (20-30% steady gains), not chip buyers. Full analysis below. For almost a decade, the most important essay in AI has been Rich Sutton’s Bitter Lesson. And for years, I…
saved by
related reading
- Does the Bitter Lesson Have Limits?dbreunig.com
- Scaling Laws, Carefully | Lil'Loglilianweng.github.io
- The Scaling Hypothesis · Gwern.netgwern.net
- Will scaling work?dwarkeshpatel.com
- The Only Important Technology Is The Internet - Kevin Lukevinlu.ai
- Will scaling work? - by Dwarkesh Patel - Dwarkesh Podcastdwarkesh.com
- Dario Amodei — "We are near the end of the exponential"dwarkesh.com
- The Bitter Lessonartfintel.com
- New Scaling Laws for Large Language Models — LessWronglesswrong.com
- The Short Case for Nvidia Stock | YouTube Transcript Optimizeryoutubetranscriptoptimizer.com
- Dario Amodei — "We are near the end of the exponential"dwarkesh.com
- There Are No New Ideas in AI… Only New Datasetsblog.jxmo.io