Claude Code Opus 4.7 Performance Tracker | Marginlab
marginlab.ai · 787 words · saved by 1 readers
Track Claude Code's daily performance on SWE-Bench-Pro. Monitor for degradation with statistical significance testing.
Claude Code Opus 4.8 Performance Tracker | Marginlab Last updated: Jun 18, 2026 Claude Code Opus 4.8 Performance Tracker We are collecting a fresh Opus 4.8 baseline on SWE tasks before resuming statistical degradation detection. • Updated daily: Daily benchmarks on a curated subset of SWE-Bench-Pro • Collect baseline: New model baseline collection is active • What you see is what you get: We benchmark in Claude Code CLI with the SOTA model (currently Opus 4.8) directly, no custom harnesses. Receive Alerts About Powered by MARGIN EVALS New model - collecting baseline data. Degradation detection
saved by
related reading
- Best practices for Claude Code - Claude Code Docsanthropic.com
- Claude Statusstatus.claude.com
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- Claude Code Cheat Sheetcc.storyfox.cz
- FrontierSWEfrontierswe.com
- A Guide to Claude Code 2.0 and getting better at using coding agents – sankalp's blogsankalp.bearblog.dev
- GitHub - ccusage/ccusage: npx ccusagegithub.com
- Parsed | Custom, interpretable AI systems that continuously learnparsed.com
- Prompting best practicesdocs.anthropic.com
- Localmaxxing - Local LLM Inference Speed Testslocalmaxxing.com
- Datacurve | The data engine for frontier AIdatacurve.ai
- Claude SWE-Bench Performance \ Anthropicanthropic.com