Beat GPT-4o at Python by searching with 100 dumb LLaMAs | Modal Blog
One thing that should be learned from the bitter lesson is the great power of general purpose methods, of methods that continue to scale with increased computation even as the available computation becomes very great. The two methods that seem to scale arbitrarily in this way are search and learning. Richard Sutton, The Bitter Lesson The eponymously distasteful take-away of Richard Sutton’s essay has often been misconstrued: because scale is all you need, they say, smaller models are doomed to irrelevance. The rapid increase in model size above one trillion parameters and the technological limitations of GPU memory together seemed to foreclose on economical frontier intelligence anywhere except at an oligopoly of intelligence-as-a-service providers. Open models and self-serve inference were in retreat. But as the quote above indicates, there are in fact two arrows in the scaling quiver: learning and search. Learning, as we do it now with neural networks, scales with memory at inference
All posts Back Research August 5, 2024 • 10 minute read Beat GPT-4o at Python by searching with 100 dumb LLaMAs Charles Frye AI Engineer Howard Halim Software Engineer View on GitHub One thing that should be learned from the bitter lesson is the great power of general purpose methods, of methods that continue to scale with increased computation even as the available computation becomes very great. The two methods that seem to scale arbitrarily in this way are search and learning . Richard Sutton, The Bitter Lesson The eponymously distasteful take-away of Richard Sutton’s essay has often been m
Explore this link on the map →related reading
- gpt-4.pdfcdn.openai.com
- The Scaling Hypothesis · Gwern.netgwern.net
- GPT-4openai.com
- I. From GPT-4 to AGI: Counting the OOMs - SITUATIONAL AWARENESSsituational-awareness.ai
- Composer2.pdfcursor.com
- Pathways Language Model (PaLM): Scaling to 540 Billion Parameters for Breakthrouai.googleblog.com
- AI progress is about to speed up | Epoch AIepoch.ai
- Will scaling work? - by Dwarkesh Patel - Dwarkesh Podcastdwarkesh.com
- What will GPT-2030 look like? — AI Alignment Forumalignmentforum.org
- The bitter lesson of LLM evalsparsed.com
- Things we learned about LLMs in 2024simonwillison.net
- o3 — LessWronglesswrong.com