Open Athena | Scaling Laws That Extrapolate 300× Past the Fit
Delphi is an open scaling suite ranging from 3e18 to 1e23 FLOPs. A pre-registered forecast from its scaling law predicted the loss of the largest run within 0.2%, extrapolating 300× past the...
In this post, I'll describe the process of developing Delphi, the Marin team's first open scaling suite, inspired by Pythia. Delphi has three parts: a scaling recipe that maps compute budgets to model configurations, a scaling suite trained from that recipe on the Google TPU Research Cloud, and a scaling law that uses the smaller Delphi models to predict the larger ones. We release the checkpoints, training mixture, and recipe so Delphi can serve as a new resource for scaling studies. A pre-registered forecast from the scaling law predicted the final loss of the largest Delphi run (1e23…
saved by
related reading
- Scaling Laws, Carefully | Lil'Loglilianweng.github.io
- [Hero Run] 535B-A23B on 18T tokensgithub.com
- The Scaling Hypothesis · Gwern.netgwern.net
- How To Scale Your Modeljax-ml.github.io
- [2509.14786] Pre-training under infinite computearxiv.org
- Fermi estimate of future training runsdanieldewey.net
- 2404.10102v1arxiv.org
- On neural scaling and the quanta hypothesisericjmichaud.com
- Demystify Transformers: A Guide to Scaling Laws | by Yu-Cheng Tsai | Sage Ai | Mediummedium.com
- IsoFLOP curves of large language models are extremely flatseverelytheoretical.wordpress.com
- New Scaling Laws for Large Language Models — LessWronglesswrong.com
- Marin 535B-A23B launch noteopenathena.ai