Answer.AI - You can now train a 70b language model at home
Today, we’re releasing Answer.AI’s first project: a fully open source system that, for the first time, can efficiently train a 70b large language model on a regular desktop computer with two or more standard gaming GPUs (RTX 3090 or 4090). This system, which combines FSDP and QLoRA, is the result of a collaboration between Answer.AI, Tim Dettmers (U Washington), and Hugging Face’s Titus von Koeller and Sourab Mangrulkar. This system will help the open source community release better models. Teknium, the creator of the extremely popular OpenHermes models and datasets, with over half a million downloads, said: “With this capability we can take huge models to new heights locally, and gigantic, hundreds of billions of parameter models are now accessible by small labs.” At Answer.AI we made this our first project because it’s a key foundation of our north star: helping make useful AI available to everyone. Just being able to use other people’s models is not enough. We want everyone to be ab
You can now train a 70b language model at home – Answer.AI Summary Today, we’re releasing Answer.AI’s first project: a fully open source system that, for the first time, can efficiently train a 70b large language model on a regular desktop computer with two or more standard gaming GPUs (RTX 3090 or 4090). This system, which combines FSDP and QLoRA, is the result of a collaboration between Answer.AI, Tim Dettmers (U Washington), and Hugging Face’s Titus von Koeller and Sourab Mangrulkar. This system will help the open source community release better models. Teknium, the creator of the extremely
Explore this link on the map →related reading
- LoRA Without Regret - Thinking Machines Labthinkingmachines.ai
- How To Scale Your Modeljax-ml.github.io
- Finetune LLMs on your own consumer hardware using tools from PyTorch and Hugging Face ecosystem – PyTorchpytorch.org
- Mosaic LLMs: GPT-3 quality formosaicml.com
- Fully Sharded Data Parallel: faster AI training with fewer GPUs Engineering at Meta -engineering.fb.com
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- Training great LLMs entirely from ground up in the wilderness as a startup - Yi Tayyitay.net
- DeepSpeed ZeRO++: A leap in speed for LLM and chat model training with 4X less communication - Microsoft Researchmicrosoft.com
- The Short Case for Nvidia Stock | YouTube Transcript Optimizeryoutubetranscriptoptimizer.com
- PiTorch: ML on Baremetal Raspberry Pis | projectsmasonjwang.com
- [1909.08053] Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelismarxiv.org
- Everything about Distributed Training and Efficient Finetuning | Sumanth's Personal Websitesumanthrh.com