Training LLMs with AMD MI250 GPUs and MosaicML
With the release of PyTorch 2.0 and ROCm 5.4, we are excited to announce that LLM training works out of the box on AMD datacenter GPUs, with zero code changes, and at high performance (144 TFLOP/s/GPU)! We are thrilled to see promising alternative options for AI hardware, and look forward to evaluating future devices and larger clusters soon.
Training LLMs with AMD MI250 GPUs and MosaicML | Databricks Blog Skip to main content With the release of PyTorch 2.0 and ROCm 5.4, we are excited to announce that LLM training works out of the box on AMD MI250 accelerators with zero code changes and at high performance! With MosaicML, the AI community has additional hardware + software options to choose from. At MosaicML, we've searched high and low for new ML training hardware on behalf of our customers. We do this to increase compute availability (as the world is in an NVIDIA supply crunch!), expand and educate the market, and ultimately re
Explore this link on the map →saved by
related reading
- Mosaic LLMs: GPT-3 quality formosaicml.com
- The Best GPUs for Deep Learning in 2023 — An In-depth Analysistimdettmers.com
- How To Scale Your Modeljax-ml.github.io
- ⭐️ Fast LLM Inference From Scratchandrewkchan.dev
- LLM Inference Performance Engineering: Best Practices | Databricks Blogdatabricks.com
- 2025: The year in LLMssimonwillison.net
- Look Ma, No Bubbles! Designing a Low-Latency Megakernel for Llama-1B · Hazy Researchhazyresearch.stanford.edu
- Training great LLMs entirely from ground up in the wilderness as a startup - Yi Tayyitay.net
- PiTorch: ML on Baremetal Raspberry Pis | projectsmasonjwang.com
- The Short Case for Nvidia Stock | YouTube Transcript Optimizeryoutubetranscriptoptimizer.com
- The Best GPUs for Deep Learning in 2023 — An In-depth Analysistimdettmers.com
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com