Latency Scaling Differences for GPT and Claude Models | Epoch AI
Epoch AI measures long-context latency scaling across four frontier models. GPT-5.6 Terra and Sol show quadratic time-to-first-token growth with context length, while Claude Sonnet 5 and Opus 5 stay near-linear.
Introduction OpenAI and Anthropic price long prompts very differently. Both have similar context windows: GPT-5.6 models have a maximum of 1.05 million tokens, while Claude models have a maximum of 1 million tokens. However, requests to OpenAI models with more than 272,000 input tokens are charged at double the normal input and cached-input price and 1.5 times the normal output price. Claude models, on the other hand, keep a fixed price regardless of input length. While pricing does not necessarily reflect the underlying inference cost, these pricing differences motivated us to test whether…
saved by
related reading
- Recursive Language Models | Alex L. Zhangalexzhang13.github.io
- My picture of the present in AI — LessWronglesswrong.com
- Late Takes on OpenAI o1alexirpan.com
- Small Models Have Arrivedcalv.info
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- GLM-5.2: Built for Long-Horizon Tasksz.ai
- Two different tricks for fast LLM inferenceseangoedecke.com
- LLM Inference Performance Engineering: Best Practices | Databricks Blogdatabricks.com
- The Oracle and the Firmcalv.info
- AINews | AINewsnews.smol.ai
- Inference characteristics of Llama · Cursorcursor.com
- AI Timeline — Complete History of 194+ Large Language Modelsllm-timeline.com