ALE-Bench: A Benchmark for Long-Horizon Objective-Driven Algorithm Engineering | alphaXiv
ALE-Bench provides the first benchmark for long-duration, score-based algorithmic programming contests, co-developed with AtCoder to evaluate AI systems on complex optimization problems. It...
Abstract How well do AI systems perform in algorithm engineering for hard optimization problems in domains such as package-delivery routing, crew scheduling, factory production planning, and power-grid balancing? We introduce ALE-Bench, a new benchmark for evaluating AI systems on score-based algorithmic programming contests. Drawing on real tasks from the AtCoder Heuristic Contests, ALE-Bench presents optimization problems that are computationally hard and admit no known exact solution. Unlike short-duration, pass/fail coding benchmarks, ALE-Bench encourages iterative solution refinement over
Explore this link on the map →related reading
- FrontierSmith: Synthesizing Open-Ended Coding Problems at Scale | alphaXivalphaxiv.org
- Composer2.pdfcursor.com
- [2203.07814] Competition-Level Code Generation with AlphaCodearxiv.org
- [2501.01257] CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratingsarxiv.org
- Are AI benchmarks doomed? - by Anson Ho and Greg Burnhamepochai.substack.com
- Competitive programming with AlphaCode — Google DeepMinddeepmind.com
- GitHub - openai/mle-bench: MLE-bench is a benchmark for measuring how well AI agents perform at machine learning engineering · GitHubgithub.com
- [2606.05405] Agents' Last Examarxiv.org
- Center for Responsible, Decentralized Intelligence at Berkeleyrdi.berkeley.edu
- PostTrainBenchposttrainbench.com
- Quantifying infrastructure noise in agentic coding evals \ Anthropicanthropic.com
- AI excels at code competitions, struggles with real workblog.peterwildeford.com