The Coding Assistant Breakdown: More Tokens Please
substack.com · 5,539 words · saved by 1 readers
Hands On With GPT 5.5, Opus 4.7, DeepSeek V4, Why Benchmarks Are Bad, and Who’s Going To Win
Since we called out the Claude Code inflection point on February 5th, we have seen a flurry of model releases. Opus, Mythos, Codex, Gemini, DeepSeek, Kimi, Qwen, GLM, MiniMax, Composer, Muse Spark, and more. Today we will break down all of these major model releases, explain when you can vs can’t trust the benchmarks, and give our predictions for the future of the agentic coding market. First we have to highlight GPT-5.5 from OpenAI. In our view, GPT-5.5 is now materially better at some tasks than all other models. We believe that GPT-5.5 has arrived at the frontier. This is a huge change…
related reading
- AINews | AINewsnews.smol.ai
- AI in 2025: gestalt — LessWronglesswrong.com
- 2025: The year in LLMssimonwillison.net
- Composer2.pdfcursor.com
- gpt-4.pdfcdn.openai.com
- GPT-4openai.com
- Introducing GPT-5.3-Codex | OpenAIopenai.com
- A Guide to Claude Code 2.0 and getting better at using coding agents – sankalp's blogsankalp.bearblog.dev
- Shipping at Inference-Speed | Peter Steinbergersteipete.me
- My picture of the present in AI — LessWronglesswrong.com
- Introducing Claude Opus 4.5 \ Anthropicanthropic.com
- TERMINAL-BENCHtbench.ai