Terminal-Bench
tbench.ai · 122 words · saved by 2 readers
A benchmark to measure and evolve with the frontier of agent work
RANKMODELAGENT 1GPT-6 Astra(max)Codex 58.2% ± 2.8% Sep 3, 20261.5B$3.3k 2Fable 5.1(max)Claude Code 57.9% ± 3.8% Sep 1, 20262.7B$6.2k 3Opus 5(max)Claude Code 51.8% ± 3.4% Jul 24, 20266.5B$6.0k 4Fable 5(max)Claude Code 44.5% ± 3.8% Jun 9, 20263.8B$7.3k 5GLM-5.3(max)Claude Code 41.8% ± 3.2% Aug 14, 20268.7B$2.7k 6GPT-5.6 Sol(max)Codex 37.3% ± 3.8% Jun 26, 20264.4B$2.5k 7Opus 4.8(max)Claude Code 23.6% ± 3.6% May 28, 20266.4B$6.5k 8GPT-5.6 Terra(max)Codex 21.5% ± 3.3% Jun 26, 20265.7B$1.7k 9Grok 4.6(high)Grok Build 20.3% ± 3.1% Aug 12, 20264.0B$3.6k 10Gemini 3.8…
saved by
related reading
- PostTrainBenchposttrainbench.com
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- Quantifying infrastructure noise in agentic coding evals \ Anthropicanthropic.com
- AINews | AINewsnews.smol.ai
- Open-Source Agentic Inference Benchmark | InferenceXinferencex.semianalysis.com
- CAIS AI Dashboarddashboard.safe.ai
- FrontierSWEfrontierswe.com
- There's An AI For That® — The front page of AItheresanaiforthat.com
- BalatroBenchbalatrobench.com
- Claude Code Cheat Sheetcc.storyfox.cz
- Datacurve | The data engine for frontier AIdatacurve.ai
- Compare AI Models: Pricing, Context & Benchmarks | OpenRouteropenrouter.ai