flâneur

>10x More Efficient Pretraining — Magic

magic.dev · 2,152 words · saved by 2 readers

Research update on compute-efficient pretraining and scaling to trillion-parameter models.

Research update on compute-efficient pretraining and scaling to trillion-parameter models. Frontier pretraining is said to be a big-lab-only game. We don’t have 100k chips yet, so there’s only one way: algorithmic efficiency. After compounding for … a while …, our pretraining recipe is now >10x more compute-efficient than that of leading open-weight base models. We match DeepSeek V4 Pro Base using ~50x fewer FLOPs – that’s around half of GPT3’s pretraining compute, or ~$0.5M on GB200. We continued scaling 10x (~$4M) and meaningfully outperformed all publicly available open base models on…

saved by

related reading