ProgramBench
programbench.com · 4,914 words · saved by 1 readers
ProgramBench evaluates whether language models can rebuild programs from scratch.
ProgramBench ./ Program Bench Can language models rebuild programs from scratch? Given only a compiled binary and its documentation, agents must architect and implement a complete codebase that reproduces the original program's behavior. John Yang * , Kilian Lieret * , Jeffrey Ma , Parth Thakkar , Dmitrii Pedchenko , Sten Sootla , Emily McMilin , Pengcheng Yin , Rui Hou , Gabriel Synnaeve , Diyi Yang , Ofir Press Meta Superintelligence Labs • Stanford University • Harvard University Leaderboard Evaluated with mini-SWE-agent · 200 tasks · Updated May. 11, 2026 · S
saved by
related reading
- Best practices for Claude Code - Claude Code Docsanthropic.com
- GitHub - x1xhlol/system-prompts-and-models-of-ai-tools: FULL Augment Code, Claude Code, Cluely, CodeBuddy, Comet, Cursor, Devin AI, Junie, Kiro, Leap.new, Lovable, Manus, NotionAI, Orchids.app, Perplexity, Poke, Qoder, Replit, Same.dev, Trae, Traycer AI, VSCode Agent, Warp.dev, Windsurf, Xcode, Z.ai Code, Dia & v0. (And other Open Sourced) System Prompts, Internal Tools & AI Modelsgithub.com
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- Claude SWE-Bench Performance \ Anthropicanthropic.com
- GitHub · Change is constant. GitHub keeps you ahead.github.com
- FrontierSWEfrontierswe.com
- Cursor: AI coding agentcursor.com
- Claude Code Cheat Sheetcc.storyfox.cz
- The open source AI coding agentopencode.ai
- crawshaw - 2025-01-06crawshaw.io
- MAI-Thinking-1 | Microsoft AImicrosoft.ai
- MirrorCode: Evidence AI can already do some weeks-long coding tasks | Epoch AIepoch.ai