FrontierCode 1.1 | Cognition
cognition.com · 1,341 words · saved by 2 readers
FrontierCode 1.1 refines our methodology to distinguish legitimate internet use from unfair use.
FrontierCode Leaderboard Benchmarks for how well models meet the standards of high-quality production codebases View Now → One month ago, we introduced FrontierCode 1 , an eval designed to measure not just code correctness, but also code quality. Today, we are releasing a refined version, FrontierCode 1.1 , with the following improvements: Fair internet use. We refined our methodology to capture the nuance between legitimate internet use (e.g., looking up documentation) and unfair use (anything that could reveal a task's solution). Fairer grading. We audited all 1000+ grading criteria and rela
saved by
related reading
- Introducing FrontierCode | Cognitioncognition.com
- FrontierSWEfrontierswe.com
- Introducing FrontierCode | Cognitioncognition.ai
- GitHub - METR/RE-Bench · GitHubgithub.com
- Building Auto Mode for Open Models — Benjamin Andersonbenanderson.work
- Senior SWE-Benchsenior-swe-bench.snorkel.ai
- FrontierSWEfrontierswe.com
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- Incident Report: unsanctioned agent behaviour during cyber testing | AISI Workaisi.gov.uk
- Recent Frontier Models Are Reward Hacking - METRmetr.org
- Frontier Risk Report (February to March 2026) - METRmetr.org
- Best practices for Claude Code - Claude Code Docsanthropic.com