flâneur

Marin 535B-A23B launch note | Open Athena

openathena.ai · 1,814 words · saved by 1 readers

A look back at Marin's first year at Open Athena and forward at the 535B-A23B hero run, the largest model the team has ever trained.

A couple weeks ago, we kicked off our capstone run of the year, a mixture of experts (MoE) model with 535B total parameters and 23B active. The plan is for it to pretrain for about 3 months on 18 trillion tokens, including agentic, coding, and scientific data. Consistent with our commitment to open development, the entire run—including architecture, data, predicted loss, etc.—is documented on GitHub. It's the largest model we've ever trained, and so far, it's going almost boringly well, lining up with our predictions and chugging along with a minimum of infrastructural or numerical fuss.…

saved by

related reading