Marin 535B-A23B launch note | Open Athena
openathena.ai · 1,814 words · saved by 1 readers
A look back at Marin's first year at Open Athena and forward at the 535B-A23B hero run, the largest model the team has ever trained.
A couple weeks ago, we kicked off our capstone run of the year, a mixture of experts (MoE) model with 535B total parameters and 23B active. The plan is for it to pretrain for about 3 months on 18 trillion tokens, including agentic, coding, and scientific data. Consistent with our commitment to open development, the entire run—including architecture, data, predicted loss, etc.—is documented on GitHub. It's the largest model we've ever trained, and so far, it's going almost boringly well, lining up with our predictions and chugging along with a minimum of infrastructural or numerical fuss.…
saved by
related reading
- Marin is charting the way to the open frontier of artificial intelligencemarin.community
- Introducing Marin: An Open Lab for Building Foundation Models | Marinmarin.community
- My picture of the present in AI — LessWronglesswrong.com
- Scaling Laws That Extrapolate 300× Past the Fitopenathena.ai
- AINews | AINewsnews.smol.ai
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- Together AI | The AI Native Cloudtogether.ai
- laguna-m1-xs2-technical-report.pdfpoolside.ai
- Fermi estimate of future training runsdanieldewey.net
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- [Hero Run] 535B-A23B on 18T tokensgithub.com
- Pathways Language Model (PaLM): Scaling to 540 Billion Parameters for Breakthrouai.googleblog.com