✳flâneur — a map of the web's best reading
Senior SWE-Bench
senior-swe-bench.snorkel.ai · 1,679 words · saved by 1 readers
Evaluating agents as senior engineers on the work we actually give them
Senior SWE-Bench Senior SWE-Bench We treat agents like senior engineers, so why evaluate them like junior engineers? 01 Senior engineers build features without over-specified requirements Senior SWE-Bench feature tasks have realistic instructions that read like natural language messages rather than over-specified requirements. To reliably evaluate these tasks, we introduce a validation agent which uses expert-designed recipes to write behavioral tests that adapt to submitted solutions. 02 Senior engineers solve bugs that require runtime investigation from behavioral reports Senior SWE-Bench bu
Explore this link on the map →saved by
related reading
- Senior SWE-Benchsenior-swe-bench.snorkel.ai
- Cookbookcookbook.openai.com
- Claude SWE-Bench Performance \ Anthropicanthropic.com
- Demystifying evals for AI agents \ Anthropicanthropic.com
- GitHub - openai/mle-bench: MLE-bench is a benchmark for measuring how well AI agents perform at machine learning engineering · GitHubgithub.com
- Building Effective AI Agents \ Anthropicanthropic.com
- AI Agent Benchmark for Real-World Professional Workflowsagents-last-exam.org
- 2502.18449arxiv.org
- Building reliable AI agents · parth sareenparthsareen.com
- [2410.06992] SWE-Bench+: Enhanced Coding Benchmark for LLMsar5iv.labs.arxiv.org
- After Automation | Everyevery.to
- Introducing SWE-grep and SWE-grep-mini: RL for Multi-Turn, Fast Context Retrieval | Cognitioncognition.ai