✳flâneur — a map of the web's best reading
Designing AI resistant technical evaluations \ Anthropic
anthropic.com · 2,674 words · saved by 1 readers
What we learned from three iterations of a performance engineering take-home that Claude keeps beating.
Written by Tristan Hume, a lead on Anthropic's performance optimization team. Tristan designed—and redesigned—the take-home test that's helped Anthropic hire dozens of performance engineers. Evaluating technical candidates becomes harder as AI capabilities improve. A take-home that distinguishes well between human skill levels today may be trivially solved by models tomorrow—rendering it useless for evaluation. Since early 2024, our performance engineering team has used a take-home test where candidates optimize code for a simulated accelerator. Over 1,000 candidates have completed it, and doz
Explore this link on the map →related reading
- How We Hire Engineers When AI Writes Our Code | Tolans.comtolans.com
- When AI builds itself \ Anthropicanthropic.com
- Demystifying evals for AI agents \ Anthropicanthropic.com
- Composer2.pdfcursor.com
- A Guide to Claude Code 2.0 and getting better at using coding agents – sankalp's blogsankalp.bearblog.dev
- How Claude Code is built - by Gergely Orosznewsletter.pragmaticengineer.com
- Humans Still Beat AI in the Long Horizon: Revisiting Test-Time Scaling in the Agent Era | Qiuyang Mangjoyemang33.github.io
- ALE-Bench: A Benchmark for Long-Horizon Objective-Driven Algorithm Engineering | alphaXivalphaxiv.org
- Introducing Claude Opus 4.5 \ Anthropicanthropic.com
- Quantifying infrastructure noise in agentic coding evals \ Anthropicanthropic.com
- Best practices for Claude Code - Claude Code Docsanthropic.com
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai