AI Benchmark Leaderboards & Model Evals | BenchmarkList
29 tracked: First-Person Fairness E... 0.61% | Cybersecurity Safeguard... 0.986 | Dynamic Mental Health B... 0.989 | +26 36 tracked: AAV Capsid Packaging Pr... 0.529 | Capture-the-Flags Chall... 96.67% | CyberGym Jailbreak Safe... 0% | +33 29 tracked: Cybersecurity Safeguard... 0.987 | MLE-Bench Revised ~71% | NanoGPT (OpenAI internal) ~14.5% | +26 21 tracked: Age of LLM: A Strategic... 1.67 | Opus Magnum Bench 9.1419 reward / 36; 15/36 solved | AIME 2026 99.2% | +18 9 tracked: Age of LLM: A Strategic... 0.70 | ObviousBench 95.83 | Opus Magnum Bench 3.9280 reward / 36; 8/36 solved | +6 72 tracked: AA-Briefcase 1586 Elo | ArxivMath 78.6% | Toolathlon 61.7% | +69 106 tracked: BioMysteryBench 84.0% | Organic chemistry (inte... 90.0% | ProgramBench (Anthropic... 93.0% | +103 16 tracked: WideSearch 62.0 | WildClawBench 47.7 | SWE Atlas - Codebase QnA 31.5 | +13 16 tracked: WideSearch 75.6 | SWE Atlas - Test Writing 40.0 | WildClawBench 53.5 | +13 1 tracked: FrontierCode 5.5% 37 tracked: Cod
New BenchmarksAll Indirect prompt-injection robustness evaluation using Gray Swan's combined attack set in transfer-only mode without computer use. Sep 1CWE-BenchCybersecurity Agents must find and fix undisclosed vulnerabilities without breaking existing tests. A verifiable benchmark for AI agents recovering biosecurity-relevant biological function from real experimental data, biophysical assays, structures, and sequences. Aug 26MMJailBenchBenchmark To address this limitation, we introduce MMJailBench, a factorized benchmark that systematically varies and combines these factors under…
saved by
related reading
- benchmarks.bio — Agentic benchmarks on messy, real-world biological databenchmarks.bio
- PostTrainBenchposttrainbench.com
- TERMINAL-BENCHtbench.ai
- CAIS AI Dashboarddashboard.safe.ai
- AI in 2025: gestalt — LessWronglesswrong.com
- Center for Responsible, Decentralized Intelligence at Berkeleyrdi.berkeley.edu
- Frontier AI Cybersecurity Observatorycybergym.io
- FrontierSWEfrontierswe.com
- Frontier Risk Report (February to March 2026) - METRmetr.org
- Import AIjack-clark.net
- Are AI benchmarks doomed? - by Anson Ho and Greg Burnhamepochai.substack.com
- GitHub - harbor-framework/terminal-bench-science: Terminal-Bench-Science: Evaluating AI agents on research workflows across scientific domainsgithub.com