God is hungry for Context: First thoughts on o3 pro
As “leaked”, OpenAI cut o3 pricing by 80% today (from $10/$40 per mtok to $2/$8 - matching GPT 4.1 pricing!!) to set the stage of the launch of o3-pro ($20/$80, supporting an unverified community theory that the -pro variants are 10x base model calls with majority voting as referenced in their papers and in our Chai episode). o3-pro reports a 64% win rate vs o3 on human testers and does marginally better on 4/4 reliability benchmarks, but as sama noticed, the actual experience expands when you test it DIFFERENTLY… I’ve had early access to o3 pro for the past week. below are my (early) thoughts: We’re in the era of task-specific models. On one hand, we have “normal” models like 3.5 Sonnet and 4o—the ones we talk to like friends, who help us with our writing, and answer our day-to-day queries. On the other, we have gigantic, slow, expensive, IQ-maxxing reasoning models that we go to for deep analysis (they’re great at criticism), one-shotting complex problems, and pushing the edge of pur
As “leaked”, OpenAI cut o3 pricing by 80% today (from $10/$40 per mtok to $2/$8 - matching GPT 4.1 pricing!!) to set the stage of the launch of o3-pro ($20/$80, supporting an unverified community theory that the -pro variants are 10x base model calls with majority voting as referenced in their papers and in our Chai episode). o3-pro reports a 64% win rate vs o3 on human testers and does marginally better on 4/4 reliability benchmarks, but as sama noticed, the actual experience expands when you test it DIFFERENTLY… Thanks to Hacker News and theo for covering us. I’ve had early access to o3…
related reading
- o3, Oh Mythezvi.substack.com
- Reverse engineering OpenAI’s o1interconnects.ai
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- Late Takes on OpenAI o1alexirpan.com
- Learning to reason with LLMs | OpenAIopenai.com
- State of AI 2025: 100T Token LLM Usage Study | OpenRouteropenrouter.ai
- o3 — LessWronglesswrong.com
- Things we learned about LLMs in 2024simonwillison.net
- AI progress is about to speed up | Epoch AIepoch.ai
- 2025: The year in LLMssimonwillison.net
- The last six months in LLMs, illustrated by pelicans on bicyclessimonwillison.net
- Latency Scaling Differences for GPT and Claude Modelsepoch.ai