FrontierSWE Blog · FrontierSWE V2
frontierswe.com · 4,619 words · saved by 2 readers
Benchmarking software engineering skill at the edge of human ability
New Tasks FrontierSWE v2 comes with 21 new challenges, bringing the total number of tasks to 34. When selecting problems, we specifically targeted new skills and domains that were not yet present in the prior version of the benchmark. Visual Reasoning Several new tasks in FrontierSWE benefit from strong visual understanding. In the ML research domain, models benefit from vision capabilities to understand training data: Snooker Prediction asks agents to build a computer vision pipeline that predicts the positions of balls in an animated snooker table video, and Vision-only TORCS Racing Bot…
saved by
related reading
- GitHub - METR/RE-Bench · GitHubgithub.com
- FrontierSWEfrontierswe.com
- Introducing FrontierCode | Cognitioncognition.com
- Composer2.pdfcursor.com
- As Rocks May Think | Eric Jangevjang.com
- Senior SWE-Benchsenior-swe-bench.snorkel.ai
- Claude SWE-Bench Performance \ Anthropicanthropic.com
- Evaluating frontier AI R&D capabilities of language model agents against human experts - METRmetr.org
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- Scaling Agentic RL: 365,000+ Environments for SWE, Terminal, and Searchprimeintellect.ai
- PostTrainBenchposttrainbench.com
- SWE-1.7: Frontier Intelligence at a Fraction of the Costcognition.com