why don’t we explicitly train models to be good at generalization?
this corresponds to why, for example, it’s better to learn from imitation in the early parts of your life — whereas you can start to learn from RL (trying things yourself, ‘FAFO’) when you’re doing quite well Filtering research ideas: Write them down in a big list of one-off ideas and cluster them over time. If the same idea keeps coming up in different forms, there’s probably something larger interesting there! - William Merrill By elimination. For example, you can rank them by expected impact, novelty, feasibility given the time frame and res… Electroporation is a physical DNA delivery method that applies a high-voltage electric field to temporarily perforate the cell membrane and drive the entry of foreign genetic material rich metaphor Finally, we demonstrate an active learning process for electroporation parameter exploration and optimization enabled by an automated robotic… There are many fundamental bottlenecks that slow biology research down. Many of them are so basic that we r
NB: Since GPT-3 showed impressive few-shot generalization abilities, we’ve backed down from calling “meta-learning” by this name explicitly and started studying it in many more ways. I’d like to write about that transition on a future day. Today, I’ll elide a lot of that detail. I do think it may be helpful to think in explicit meta-learning terms to solve RL sample-inefficiency. Meta-RL seems like the obvious solution to the sample-inefficiency of RL. Meta-learning (or its special case, “in-context learning”) remains the fastest route to few-shot generalization. Despite that, this year’s…
saved by
related reading
- Just Ask for Generalization | Eric Jangevjang.com
- How can LLM RL Work Despite Information-Theoretic Inefficiencyberen.io
- The Scaling Hypothesis · Gwern.netgwern.net
- The Extreme Inefficiency of RL for Frontier Models - Toby Ordtobyord.com
- The upcoming GPT-3 moment for RL | Mechanize, Inc.mechanize.work
- Meta-Learning: Learning to Learn Fastlilianweng.github.io
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- RL is even more information inefficient than you thoughtdwarkesh.com
- Evolution as Backstop for Reinforcement Learning · Gwern.netgwern.net
- Ilya Sutskever — We're moving from the age of scaling to the age of researchdwarkesh.com
- State of RL for reasoning LLMs | A. Weersaweers.de
- To Understand Language is to Understand Generalization | Eric Jangevjang.com