Headroom for AI development – Machine Learning (Theory)
In thinking about what are good research problems, it’s sometimes helpful to switch from what is understood to what is clearly possible. This encourages us to think beyond simply improving the existing system. For example, we have seen instances throughout the history of machine learning where researchers have argued for fixing an architecture and using it for short-term success, ignoring potential for long-term disruption. As an example, the speech recognition community spent decades focusing on Hidden Markov Models at the expense of other architectures, before eventually being disrupted by advancements in deep learning. Support Vector Machines were disrupted by deep learning, and convolutional neural networks were displaced by transformers. This pattern may repeat for the current transformer/large language model (LLM) paradigm. Here are some quick calculations suggesting it may be possible to do significantly better along multiple axes. Examples include the following: The core of thi
( Dylan Foster and Alex Lamb both helped in creating this.) In thinking about what are good research problems, it’s sometimes helpful to switch from what is understood to what is clearly possible. This encourages us to think beyond simply improving the existing system. For example, we have seen instances throughout the history of machine learning where researchers have argued for fixing an architecture and using it for short-term success, ignoring potential for long-term disruption. As an example, the speech recognition community spent decades focusing on Hidden Markov Models at the expense of
Explore this link on the map →saved by
related reading
- GenAI Handbookgenai-handbook.github.io
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- Transformer Circuits Threadtransformer-circuits.pub
- On the Tradeoffs of SSMs and Transformers | Goomba Labgoombalab.github.io
- I am worried about near-term non-LLM AI developments — LessWronglesswrong.com
- There Are No New Ideas in AI… Only New Datasetsblog.jxmo.io
- I. From GPT-4 to AGI: Counting the OOMs - SITUATIONAL AWARENESSsituational-awareness.ai
- A Short History of Artificial Intelligenceevery.to
- Large Language Models Reading List | Sebastian Raschka, PhDsebastianraschka.com
- What I've Learned About AI in the Past Two Months.sheracaolity.ghost.io
- Memory makes computation universal, remember?thinks.lol
- Norman Mu | The Myth of Data Inefficiency in Large Language Modelsnormanmu.com