✳flâneur — a map of the web's best reading
ProgramBench
programbench.com · 4,914 words · saved by 1 readers
ProgramBench evaluates whether language models can rebuild programs from scratch.
ProgramBench ./ Program Bench Can language models rebuild programs from scratch? Given only a compiled binary and its documentation, agents must architect and implement a complete codebase that reproduces the original program's behavior. John Yang * , Kilian Lieret * , Jeffrey Ma , Parth Thakkar , Dmitrii Pedchenko , Sten Sootla , Emily McMilin , Pengcheng Yin , Rui Hou , Gabriel Synnaeve , Diyi Yang , Ofir Press Meta Superintelligence Labs • Stanford University • Harvard University Leaderboard Evaluated with mini-SWE-agent · 200 tasks · Updated May. 11, 2026 · S
Explore this link on the map →saved by
related reading
- Best practices for Claude Code - Claude Code Docsanthropic.com
- Composer2.pdfcursor.com
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- Claude SWE-Bench Performance \ Anthropicanthropic.com
- Claude Code Cheat Sheetcc.storyfox.cz
- MirrorCode: Evidence AI can already do some weeks-long coding tasks | Epoch AIepoch.ai
- GitHub - shroominic/codeinterpreter-api: 👾 Open source implementation of the ChatGPT Code Interpreter · GitHubgithub.com
- GitHub - oughtinc/ice: Interactive Composition Explorer: a debugger for compositional language model programs · GitHubgithub.com
- Codex use casesdevelopers.openai.com
- crawshaw - 2025-01-06crawshaw.io
- GitHub - x1xhlol/system-prompts-and-models-of-ai-tools: FULL Augment Code, Claude Code, Cluely, CodeBuddy, Comet, Cursor, Devin AI, Junie, Kiro, Leap.new, Lovable, Manus, NotionAI, Orchids.app, Perplexity, Poke, Qoder, Replit, Same.dev, Tragithub.com
- Claude Code Opus 4.8 Performance Tracker | Marginlabmarginlab.ai