ChatGPT 5.1 Codex Max - by Zvi Mowshowitz
It scores 77.9% on SWE-bench-verified, 79.9% on SWE-Lancer-IC SWE and 58.1% on Terminal-Bench 2.0, all substantial gains over GPT-5.1-Codex. It’s triggering OpenAI to prepare for being high level in cybersecurity threats. There’s a 27 page system card. One could call this the secret ‘real’ GPT-5.1 that matters. They even finally trained it to use Windows, somehow this is a new idea. My goal is for my review of Opus 4.5 to start on Friday, as it takes a few days to sort through new releases. This post was written before Anthropic revealed Opus 4.5, and we don’t yet know how big an upgrade Opus 4.5 will prove to be. As always, try all your various options and choose what is best for you. GPT-5.1-Codex-Max is a new high on the METR graph. METR’s thread is here. Prinz: METR (50% accuracy): GPT-5.1-Codex-Max = 2 hours, 42 minutes This is 25 minutes longer than GPT-5. Samuel Albanie: a data point for that ai 2027 graph That’s in between the two lines, looking closer to linear progress. Fin
OpenAI has given us GPT-5.1-Codex-Max, their best coding model for OpenAI Codex. They claim it is faster, more capable and token-efficient and has better persistence on long tasks. It scores 77.9% on SWE-bench-verified, 79.9% on SWE-Lancer-IC SWE and 58.1% on Terminal-Bench 2.0, all substantial gains over GPT-5.1-Codex. It’s triggering OpenAI to prepare for being high level in cybersecurity threats. There’s a 27 page system card. One could call this the secret ‘real’ GPT-5.1 that matters. They even finally trained it to use Windows, somehow this is a new idea. My goal is for my review…
related reading
- Introducing GPT-5.3-Codex | OpenAIopenai.com
- The Coding Assistant Breakdown: More Tokens Pleasesubstack.com
- gpt-4.pdfcdn.openai.com
- GPT-4openai.com
- Shipping at Inference-Speed | Peter Steinbergersteipete.me
- GPT-5.6 Preview System Carddeploymentsafety.openai.com
- AINews | AINewsnews.smol.ai
- GPT-5.6 Preview System Card - OpenAI Deployment Safety Hubdeploymentsafety.openai.com
- GPT-5.5 and the broken state of government evalstransformernews.ai
- Introducing Claude Fable 5.1 and Claude Mythos 5.1anthropic.com
- GLM-5.1: Towards Long-Horizon Tasksz.ai
- AI in 2025: gestalt — LessWronglesswrong.com