How many days did it take to train GPT-3? Is training a neural net model a parallelizable task? : GPT3
How many days did it take to train the GPT-3 model? From the above table it says that it took 3640 days of training for GPT-3. That is 9.97 years. Am I right? If then how did they train the model for a company that was setup 5 years ago? Is training a neural net model a parallelizable task for them to train on many GPUs in parallel and reduce the time needed to train? In my opinion training aka optimising the weights cannot be a parallelizable task as each weight have to be optimised step by step slowly through each back-propagation. Each weight will reach the optimum value only by changing it's value little by little in sequential order. So it cannot be a parrallelizable task. Am I right? What does tokens mean in this table? Great questions... It would take 355 years to train GPT-3 on a single NVIDIA Tesla V100 GPU. OpenAI launched GPT-3 in May/2020. Microsoft (using Azure DCs) built a supercomputer with 10,000 V100 GPUs exclusively for OpenAI. Estimated that it cost around $5M in com
Reddit - Please wait for verification
Explore this link on the map →related reading
- How To Scale Your Modeljax-ml.github.io
- The Scaling Hypothesis · Gwern.netgwern.net
- What Is ChatGPT Doing … and Why Does It Work?-Stephen Wolfram Writingswritings.stephenwolfram.com
- What will GPT-2030 look like? — AI Alignment Forumalignmentforum.org
- The Short Case for Nvidia Stock | YouTube Transcript Optimizeryoutubetranscriptoptimizer.com
- Transformer Math 101 | EleutherAI Blogblog.eleuther.ai
- microgptkarpathy.github.io
- Training great LLMs entirely from ground up in the wilderness as a startup - Yi Tayyitay.net
- Estimating training compute of deep learning models | Epoch AIepochai.org
- Pipeline-Parallelism: Distributed Training via Model Partitioningsiboehm.com
- Fermi estimate of future training runsdanieldewey.net
- The Best GPUs for Deep Learning in 2023 — An In-depth Analysistimdettmers.com