Training great LLMs entirely from ground zero in the wilderness as a startup — Yi Tay
Given that we’ve successfully trained pretty strong multimodal language models at Reka, many people have been particularly curious about the experiences of building infrastructure and training large language & multimodal models from scratch from a completely clean slate. I complain a lot about external (outside Google) infrastructure and code on my social media, leading people to really be curious about what are the things I miss and what I hate/love in the wilderness. So here’s a post (finally). This blogpost sheds light on the challenges and lessons learned. I hope this post will be interesting and/or educational for many. Training LLMs in the wilderness (image generated by Dall-E) The first requisite for training models is acquiring compute. This seems straightforward and easy enough. However, the largest surprise turned out to be the instability of compute providers and how large variance the quality of clusters, accelerators and their connectivity were depending on the source. Peo
Training great LLMs entirely from ground up in the wilderness as a startup Mar 6 Written By Yi Tay Given that we’ve successfully trained pretty strong multimodal language models at Reka, many people have been particularly curious about the experiences of building infrastructure and training large language & multimodal models from scratch from a completely clean slate. I complain a lot about external (outside Google) infrastructure and code on my social media, leading people to really be curious about what are the things I miss and what I hate/love in the wilderness. So here’s a post (finally).
Explore this link on the map →saved by
related reading
- How To Scale Your Modeljax-ml.github.io
- Mosaic LLMs: GPT-3 quality formosaicml.com
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- Reinforcement learning is an infrastructure problemmodal.com
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- Multi-Datacenter Training: OpenAI's Ambitious Plan To Beat Google's Infrastructuresemianalysis.com
- The Llama Hitchiking Guide to Local LLMs – hackerllamaosanseviero.github.io
- Composer2.pdfcursor.com
- From bare metal to a 70B model: infrastructure set-up and scripts - Imbueimbue.com
- the world’s largest distributed LLM training job on TPU v5e | Google Cloud Blogcloud.google.com
- Training LLMs with AMD MI250 GPUs and MosaicML | Databricks Blogmosaicml.com
- A Hitchhiker’s Guide to ML Training Infrastructure | CMU Software Engineering Institutesei.cmu.edu