flâneur — a map of the web's best reading

Beat GPT-4o at Python by searching with 100 dumb LLaMAs | Modal Blog

modal.com · 1,607 words · saved by 1 readers

One thing that should be learned from the bitter lesson is the great power of general purpose methods, of methods that continue to scale with increased computation even as the available computation becomes very great. The two methods that seem to scale arbitrarily in this way are search and learning. Richard Sutton, The Bitter Lesson The eponymously distasteful take-away of Richard Sutton’s essay has often been misconstrued: because scale is all you need, they say, smaller models are doomed to irrelevance. The rapid increase in model size above one trillion parameters and the technological limitations of GPU memory together seemed to foreclose on economical frontier intelligence anywhere except at an oligopoly of intelligence-as-a-service providers. Open models and self-serve inference were in retreat. But as the quote above indicates, there are in fact two arrows in the scaling quiver: learning and search. Learning, as we do it now with neural networks, scales with memory at inference

All posts Back Research August 5, 2024 • 10 minute read Beat GPT-4o at Python by searching with 100 dumb LLaMAs Charles Frye AI Engineer Howard Halim Software Engineer View on GitHub One thing that should be learned from the bitter lesson is the great power of general purpose methods, of methods that continue to scale with increased computation even as the available computation becomes very great. The two methods that seem to scale arbitrarily in this way are search and learning . Richard Sutton, The Bitter Lesson The eponymously distasteful take-away of Richard Sutton’s essay has often been m

Explore this link on the map →

related reading