flâneur — a map of the web's best reading

From bare metal to a 70B model: infrastructure set-up and scripts - imbue

imbue.com · 6,609 words · saved by 3 readers

We would like to thank Voltage Park, Dell, H5, and NVIDIA for their invaluable partnership and help with setting up our cluster. A special thanks to Ozan, Melissa, Drew, Michael, and David at Voltage Park for their dedicated support throughout the project. Setting up a cluster of this size is an enormous challenge and we couldn’t have done it without them. This is the second of a three-part series on how we trained our 70B model. We covered setting up infrastructure, conducting evaluations, and hyperparameter optimization. In the span of a few months, with a small team of researchers and engineers, we trained a 70B parameter model from scratch on our own infrastructure that outperformed zero-shot GPT-4o on reasoning-related tasks. Today, we’re sharing an end-to-end guide for setting up the required infrastructure: from bringing up the initial cluster and installing the OS, to automatically recovering from errors encountered during training. In each step, we detail the challenges we enc

From bare metal to a 70B model: infrastructure set-up and scripts - Imbue Article / Research From bare metal to a 70B model: infrastructure set-up and scripts 46 min read Last updated 30 Apr 2026 Introduction Background: How this is supposed to work Process: How to go from bare metal to a fully operational cluster Provisioning individual machines Provisioning InfiniBand Ensuring fully healthy machines Diagnosing common training issues Training slowdowns (as measured by MFU) Improving infrastructure tooling Reflections and learnings Conclusion We would like to thank Voltage Park, Dell, H5, and

Explore this link on the map →

saved by

related reading