Large Language Models: A New Moore's Law?
huggingface.co · 1,300 words · saved by 1 readers
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
A few days ago, Microsoft and NVIDIA introduced Megatron-Turing NLG 530B, a Transformer-based model hailed as "the world’s largest and most powerful generative language model." This is an impressive show of Machine Learning engineering, no doubt about it. Yet, should we be excited about this mega-model trend? I, for one, am not. Here's why. This is your Brain on Deep Learning Researchers estimate that the human brain contains an average of 86 billion neurons and 100 trillion synapses. It's safe to assume that not all of them are dedicated to language either. Interestingly, GPT-4 is…
related reading
- The Scaling Hypothesis · Gwern.netgwern.net
- Fermi estimate of future training runsdanieldewey.net
- The Short Case for Nvidia Stock | YouTube Transcript Optimizeryoutubetranscriptoptimizer.com
- Scaling: The State of Play in AIoneusefulthing.org
- Why didn't we get GPT-2 in 2005?dynomight.net
- Chinchillaarxiv.org
- GPT-4 Architecture, Infrastructure, Training Dataset, Costs, Vision, MoEsemianalysis.com
- The Little Book of Deep Learningfleuret.org
- What will GPT-2030 look like? — AI Alignment Forumalignmentforum.org
- New Scaling Laws for Large Language Models — LessWronglesswrong.com
- Large Language Models Reading List | Sebastian Raschka, PhDsebastianraschka.com
- Navigating the High Cost of AI Compute | Andreessen Horowitza16z.com