✳flâneur — a map of the web's best reading
FrontierCode 1.1 | Cognition
cognition.com · 1,341 words · saved by 2 readers
FrontierCode 1.1 refines our methodology to distinguish legitimate internet use from unfair use.
FrontierCode Leaderboard Benchmarks for how well models meet the standards of high-quality production codebases View Now → One month ago, we introduced FrontierCode 1 , an eval designed to measure not just code correctness, but also code quality. Today, we are releasing a refined version, FrontierCode 1.1 , with the following improvements: Fair internet use. We refined our methodology to capture the nuance between legitimate internet use (e.g., looking up documentation) and unfair use (anything that could reveal a task's solution). Fairer grading. We audited all 1000+ grading criteria and rela
Explore this link on the map →saved by
related reading
- Introducing FrontierCode | Cognitioncognition.ai
- GitHub - METR/RE-Bench · GitHubgithub.com
- Introducing FrontierCode | Cognitioncognition.com
- Demystifying evals for AI agents \ Anthropicanthropic.com
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- Composer2.pdfcursor.com
- Best practices for Claude Code - Claude Code Docsanthropic.com
- Recent Frontier Models Are Reward Hacking - METRmetr.org
- Fara1.5 - A family of frontier computer use agent models - Microsoft Researchmicrosoft.com
- Frontier Risk Report (February to March 2026) - METRmetr.org
- GitHub - ccusage/ccusage: npx ccusage · GitHubgithub.com
- Quantifying infrastructure noise in agentic coding evals \ Anthropicanthropic.com