flâneur

Optimally allocating compute between inference and training | Epoch AI

epoch.ai · 2,104 words · saved by 1 readers

AI labs should spend comparable resources on training and inference, assuming they can flexibly balance compute between the two to maintain performance.

Introduction Sam Altman recently claimed that OpenAI currently generates around 100 billion tokens per day, or about 36 trillion tokens per year. Given that modern language models are trained on the order of 10 trillion tokens and tokens seen during training are around three times more expensive1 compared to tokens seen or generated during inference, a naive analysis2 suggests OpenAI’s annual inference costs are on the same order as their annual model training costs. This seems like an odd coincidence at first: why should one of these not completely dominate the other? However, there’s a…

saved by

related reading