FrontierSWE
frontierswe.com · 192 words · saved by 3 readers
Benchmarking software engineering skill at the edge of human ability
FrontierSWEV2 Benchmarking software engineering skill at the edge of human ability. By Leaderboard Per provider mean@5worst→best@5 Scores across all 34 tasks. Each model runs 5 trials per task with a 20-hour budget. #ModelScoreAvg CostTime 1 Claude Fable 5.1 proximus 56.3 % ± 11.1 $138.5511.6h 2 GPT-5.6 proximus 32.2 % ± 11.0 $179.648.6h 3 GLM-5.3 proximus 30.2 % ± 11.5 $97.2217.0h 4 Kimi K3 proximus 25.9 % ± 11.8 $109.7118.4h 5 Grok 4.6 proximus 25.3 % ± 12.2 $243.4313.8h 6 Gemini 3.7 Flash proximus 20.3 % ± 10.5 $34.147.9h 7…
saved by
related reading
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- FrontierSWEfrontierswe.com
- Datacurve | The data engine for frontier AIdatacurve.ai
- UI Skills for Design Engineers | UI Skillsui-skills.com
- CAIS AI Dashboarddashboard.safe.ai
- Parsed | Custom, interpretable AI systems that continuously learnparsed.com
- There's An AI For That® — The front page of AItheresanaiforthat.com
- BalatroBenchbalatrobench.com
- Claude Code Opus 4.8 Performance Tracker | Marginlabmarginlab.ai
- Claude SWE-Bench Performance \ Anthropicanthropic.com
- Intelligenceintelligence.ai
- Compare AI Models: Pricing, Context & Benchmarks | OpenRouteropenrouter.ai