From bare metal to a 70B model: infrastructure set-up and scripts - imbue
We would like to thank Voltage Park, Dell, H5, and NVIDIA for their invaluable partnership and help with setting up our cluster. A special thanks to Ozan, Melissa, Drew, Michael, and David at Voltage Park for their dedicated support throughout the project. Setting up a cluster of this size is an enormous challenge and we couldn’t have done it without them. This is the second of a three-part series on how we trained our 70B model. We covered setting up infrastructure, conducting evaluations, and hyperparameter optimization. In the span of a few months, with a small team of researchers and engineers, we trained a 70B parameter model from scratch on our own infrastructure that outperformed zero-shot GPT-4o on reasoning-related tasks. Today, we’re sharing an end-to-end guide for setting up the required infrastructure: from bringing up the initial cluster and installing the OS, to automatically recovering from errors encountered during training. In each step, we detail the challenges we enc
From bare metal to a 70B model: infrastructure set-up and scripts - Imbue Article / Research From bare metal to a 70B model: infrastructure set-up and scripts 46 min read Last updated 30 Apr 2026 Introduction Background: How this is supposed to work Process: How to go from bare metal to a fully operational cluster Provisioning individual machines Provisioning InfiniBand Ensuring fully healthy machines Diagnosing common training issues Training slowdowns (as measured by MFU) Improving infrastructure tooling Reflections and learnings Conclusion We would like to thank Voltage Park, Dell, H5, and
Explore this link on the map →saved by
related reading
- Training great LLMs entirely from ground up in the wilderness as a startup - Yi Tayyitay.net
- 2410.21680arxiv.org
- Multi-Datacenter Training: OpenAI's Ambitious Plan To Beat Google's Infrastructuresemianalysis.com
- How to Rack 30 Petabytes of Storage | blogsi.inc
- ClusterMAX™ 2.0: The Industry Standard GPU Cloud Rating Systemnewsletter.semianalysis.com
- The Short Case for Nvidia Stock | YouTube Transcript Optimizeryoutubetranscriptoptimizer.com
- The Best GPUs for Deep Learning in 2023 — An In-depth Analysistimdettmers.com
- How To Scale Your Modeljax-ml.github.io
- Accelerate AI & Machine Learning Workflows | NVIDIA Run:airun.ai
- PiTorch: ML on Baremetal Raspberry Pis | projectsmasonjwang.com
- A Hitchhiker’s Guide to ML Training Infrastructure | CMU Software Engineering Institutesei.cmu.edu
- Reinforcement learning is an infrastructure problemmodal.com