Beat GPT-4o at Python by searching with 100 dumb LLaMAs | Modal Blog
One thing that should be learned from the bitter lesson is the great power of general purpose methods, of methods that continue to scale with increased computation even as the available computation becomes very great. The two methods that seem to scale arbitrarily in this way are search and learning. Richard Sutton, The Bitter Lesson The eponymously distasteful take-away of Richard Sutton’s essay has often been misconstrued: because scale is all you need, they say, smaller models are doomed to irrelevance. The rapid increase in model size above one trillion parameters and the technological limitations of GPU memory together seemed to foreclose on economical frontier intelligence anywhere except at an oligopoly of intelligence-as-a-service providers. Open models and self-serve inference were in retreat. But as the quote above indicates, there are in fact two arrows in the scaling quiver: learning and search. Learning, as we do it now with neural networks, scales with memory at inference
All posts Back Research August 5, 2024 • 10 minute read Beat GPT-4o at Python by searching with 100 dumb LLaMAs Charles Frye AI Engineer Howard Halim Software Engineer View on GitHub One thing that should be learned from the bitter lesson is the great power of general purpose methods, of methods that continue to scale with increased computation even as the available computation becomes very great. The two methods that seem to scale arbitrarily in this way are search and learning . Richard Sutton, The Bitter Lesson The eponymously distasteful take-away of Richard Sutton’s essay has often been m
related reading
- gpt-4.pdfcdn.openai.com
- The Scaling Hypothesis · Gwern.netgwern.net
- GPT-4openai.com
- As Rocks May Think | Eric Jangevjang.com
- I. From GPT-4 to AGI: Counting the OOMs - SITUATIONAL AWARENESSsituational-awareness.ai
- Scaling: The State of Play in AIoneusefulthing.org
- Pathways Language Model (PaLM): Scaling to 540 Billion Parameters for Breakthrouai.googleblog.com
- AI progress is about to speed up | Epoch AIepoch.ai
- Will scaling work? - by Dwarkesh Patel - Dwarkesh Podcastdwarkesh.com
- What will GPT-2030 look like? — AI Alignment Forumalignmentforum.org
- Things we learned about LLMs in 2024simonwillison.net
- From Apples to Strawberriestmychow.substack.com