flâneur

FrontierSWE

frontierswe.com · 192 words · saved by 3 readers

Benchmarking software engineering skill at the edge of human ability

FrontierSWEV2 Benchmarking software engineering skill at the edge of human ability. By Leaderboard Per provider mean@5worst→best@5 Scores across all 34 tasks. Each model runs 5 trials per task with a 20-hour budget. #ModelScoreAvg CostTime 1 Claude Fable 5.1 proximus 56.3 % ± 11.1 $138.5511.6h 2 GPT-5.6 proximus 32.2 % ± 11.0 $179.648.6h 3 GLM-5.3 proximus 30.2 % ± 11.5 $97.2217.0h 4 Kimi K3 proximus 25.9 % ± 11.8 $109.7118.4h 5 Grok 4.6 proximus 25.3 % ± 12.2 $243.4313.8h 6 Gemini 3.7 Flash proximus 20.3 % ± 10.5 $34.147.9h 7…

saved by

related reading