Mosaic LLMs (Part 1): Billion-Parameter GPT Training Made Easy
In Part 1 of this LLM blog post series, we use the MosaicML platform to train vanilla GPT-3 models up to 1.3B params, and show how to cut training times down to hours with strong multi-node scaling. We also discover that larger models can train more efficiently than smaller models on modern hardware, and that a 10x in parameter count may only result in ~5x the training time.
mousedown mouseup touchcancel touchend touchstart auxclick dblclick pointercancel pointerdown pointerup dragend dragstart drop compositionend compositionstart keydown keypress keyup input textInput copy cut paste click change contextmenu reset submit abort auxClick cancel canPlay canPlayThrough click close contextMenu copy cut drag dragEnd dragEnter dragExit dragLeave dragOver dragStart drop durationChange emptied encrypted ended error gotPointerCapture input invalid keyDown keyPress keyUp load loadedData loadedMetadata loadStart lostPointerCapture mouseDown mouseMove mouseOut mouseOver mouseU
Explore this link on the map →related reading
- Mosaic LLMs: GPT-3 quality formosaicml.com
- Chloe's Gardenchloeyan.me
- Training great LLMs entirely from ground up in the wilderness as a startup - Yi Tayyitay.net
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- Pathways Language Model (PaLM): Scaling to 540 Billion Parameters for Breakthrouai.googleblog.com
- Go smol or go home | Harm de Vriesharmdevries.com
- microgptkarpathy.github.io
- GitHub - karpathy/nanochat: The best ChatGPT that $100 can buy. · GitHubgithub.com
- The Ultra-Scale Playbook - a Hugging Face Space by nanotronhuggingface.co
- The Ultra-Scale Playbook - a Hugging Face Space by nanotronhuggingface.co
- Sorrel Salb | Are.naare.na
- GitHub - linkedin/Liger-Kernel: Efficient Triton Kernels for LLM Training · GitHubgithub.com