flâneur — a map of the web's best reading

Senior SWE-Bench

senior-swe-bench.snorkel.ai · 1,679 words · saved by 1 readers

Evaluating agents as senior engineers on the work we actually give them

Senior SWE-Bench Senior SWE-Bench We treat agents like senior engineers, so why evaluate them like junior engineers? 01 Senior engineers build features without over-specified requirements Senior SWE-Bench feature tasks have realistic instructions that read like natural language messages rather than over-specified requirements. To reliably evaluate these tasks, we introduce a validation agent which uses expert-designed recipes to write behavioral tests that adapt to submitted solutions. 02 Senior engineers solve bugs that require runtime investigation from behavioral reports Senior SWE-Bench bu

Explore this link on the map →

saved by

related reading