✳flâneur — a map of the web's best reading
EVMBench Leaderboard — AI Smart Contract Auditors
testmachine.ai · 798 words · saved by 1 readers
117 vulnerabilities. 40 Code4rena audits. One scoreboard.
EVMBench Leaderboard — AI Smart Contract Auditors Home How It Works Token Custody Azimuth Leaderboard Blog Contact Launch App Open Benchmark · AI Security The EVMBench Leaderboard EVMBench is a standardized benchmark built by OpenAI for AI vulnerability detection on EVM smart contracts: 117 ground-truth vulnerabilities across 40 Code4rena audits . Vendors keep publishing one-off numbers. No one has put them on a single board — so we did. 40 Repositories 117 Vulnerabilities (120 at launch) 10 Published results # Model / Agent Detection Recall Found 1 Azimuth Our entry TestMachine Combined run a
Explore this link on the map →saved by
related reading
- AuditAgentauditagent.nethermind.io
- Measuring LLMs’ ability to develop exploits \ Anthropicred.anthropic.com
- Assessing Claude Mythos Preview’s cybersecurity capabilities \ Anthropicred.anthropic.com
- Security incident disclosure — July 2026huggingface.co
- Center for Responsible, Decentralized Intelligence at Berkeleyrdi.berkeley.edu
- AI Benchmark Leaderboards & Model Evals | BenchmarkListbenchmarklist.com
- Mythos finds a curl vulnerability | daniel.haxx.sedaniel.haxx.se
- AI agents find smart contract exploits \ Anthropicred.anthropic.com
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- A shallow dive into formal verificationvitalik.eth.limo
- AuditBenchalignment.anthropic.com
- GitHub - openai/mle-bench: MLE-bench is a benchmark for measuring how well AI agents perform at machine learning engineering · GitHubgithub.com