✳flâneur — a map of the web's best reading
Claude Code Opus 4.7 Performance Tracker | Marginlab
marginlab.ai · 787 words · saved by 1 readers
Track Claude Code's daily performance on SWE-Bench-Pro. Monitor for degradation with statistical significance testing.
Claude Code Opus 4.8 Performance Tracker | Marginlab Last updated: Jun 18, 2026 Claude Code Opus 4.8 Performance Tracker We are collecting a fresh Opus 4.8 baseline on SWE tasks before resuming statistical degradation detection. • Updated daily: Daily benchmarks on a curated subset of SWE-Bench-Pro • Collect baseline: New model baseline collection is active • What you see is what you get: We benchmark in Claude Code CLI with the SOTA model (currently Opus 4.8) directly, no custom harnesses. Receive Alerts About Powered by MARGIN EVALS New model - collecting baseline data. Degradation detection
Explore this link on the map →saved by
related reading
- Best practices for Claude Code - Claude Code Docsanthropic.com
- Claude Statusstatus.claude.com
- Claude Code Cheat Sheetcc.storyfox.cz
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- A Guide to Claude Code 2.0 and getting better at using coding agents – sankalp's blogsankalp.bearblog.dev
- Perplexityperplexity.ai
- GitHub - affaan-m/ECC: The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond. · GitHubgithub.com
- GitHub - ccusage/ccusage: npx ccusage · GitHubgithub.com
- GitHub - brexhq/prompt-engineering: Tips and tricks for working with Large Language Models like OpenAI's GPT-4. · GitHubgithub.com
- Branches · HazyResearch/intelligence-per-watt · GitHubgithub.com
- Claude Code Hit Escape Velocity in January. So I Used Claude Code to /Catch-Up.zerotopete.com
- cchistory: Tracking Claude Code System Prompt and Tool Changesmariozechner.at