AI Benchmark Leaderboards & Model Evals | BenchmarkList
29 tracked: First-Person Fairness E... 0.61% | Cybersecurity Safeguard... 0.986 | Dynamic Mental Health B... 0.989 | +26 36 tracked: AAV Capsid Packaging Pr... 0.529 | Capture-the-Flags Chall... 96.67% | CyberGym Jailbreak Safe... 0% | +33 29 tracked: Cybersecurity Safeguard... 0.987 | MLE-Bench Revised ~71% | NanoGPT (OpenAI internal) ~14.5% | +26 21 tracked: Age of LLM: A Strategic... 1.67 | Opus Magnum Bench 9.1419 reward / 36; 15/36 solved | AIME 2026 99.2% | +18 9 tracked: Age of LLM: A Strategic... 0.70 | ObviousBench 95.83 | Opus Magnum Bench 3.9280 reward / 36; 8/36 solved | +6 72 tracked: AA-Briefcase 1586 Elo | ArxivMath 78.6% | Toolathlon 61.7% | +69 106 tracked: BioMysteryBench 84.0% | Organic chemistry (inte... 90.0% | ProgramBench (Anthropic... 93.0% | +103 16 tracked: WideSearch 62.0 | WildClawBench 47.7 | SWE Atlas - Codebase QnA 31.5 | +13 16 tracked: WideSearch 75.6 | SWE Atlas - Test Writing 40.0 | WildClawBench 53.5 | +13 1 tracked: FrontierCode 5.5% 37 tracked: Cod
Top Models by Experimental ECI Bar = experimental ECI score Claude Fable 5 Anthropic 157.19 Claude Mythos 5 Anthropic 155.19 Claude Mythos Preview Anthropic 153.41 Claude Opus 4.8 Anthropic 151.87 GPT-5.5 Pro OpenAI 150.73
Explore this link on the map →saved by
related reading
- benchmarks.bio — Agentic AI benchmarks on messy, real-world biological databenchmarks.bio
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- There's An AI For That® — The front page of AItheresanaiforthat.com
- Parsed | Custom, interpretable AI systems that continuously learnparsed.com
- Contra Labs - Powered by Contracontralabs.com
- GitHub - brexhq/prompt-engineering: Tips and tricks for working with Large Language Models like OpenAI's GPT-4. · GitHubgithub.com
- Branches · HazyResearch/intelligence-per-watt · GitHubgithub.com
- GitHub - SoyGema/pulling_ace · GitHubgithub.com
- GitHub - open-compass/opencompass: OpenCompass is an LLM evaluation platform, supporting a wide range of models (Llama3, Mistral, InternLM2,GPT-4,LLaMa2, Qwen,GLM, Claude, etc) over 100+ datasets. · GitHubgithub.com
- Neuronpedianeuronpedia.org
- GitHub - inverse-scaling/prize: A prize for finding tasks that cause large language models to show inverse scaling · GitHubgithub.com
- ML Contestsmlcontests.com