ALE-Bench: A Benchmark for Long-Horizon Objective-Driven Algorithm Engineering | alphaXiv
ALE-Bench provides the first benchmark for long-duration, score-based algorithmic programming contests, co-developed with AtCoder to evaluate AI systems on complex optimization problems. It...
Abstract How well do AI systems perform in algorithm engineering for hard optimization problems in domains such as package-delivery routing, crew scheduling, factory production planning, and power-grid balancing? We introduce ALE-Bench, a new benchmark for evaluating AI systems on score-based algorithmic programming contests. Drawing on real tasks from the AtCoder Heuristic Contests, ALE-Bench presents optimization problems that are computationally hard and admit no known exact solution. Unlike short-duration, pass/fail coding benchmarks, ALE-Bench encourages iterative solution refinement over
saved by
related reading
- FrontierSmith: Synthesizing Open-Ended Coding Problems at Scale | alphaXivalphaxiv.org
- EdgeBench | Scaling Laws of Environment Learningedge-bench.org
- Composer2.pdfcursor.com
- [2203.07814] Competition-Level Code Generation with AlphaCodearxiv.org
- [2501.01257] CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratingsarxiv.org
- Are AI benchmarks doomed? - by Anson Ho and Greg Burnhamepochai.substack.com
- Every Benchmark is Brokenjonathanpgabor.substack.com
- GitHub - openai/mle-bench: MLE-bench is a benchmark for measuring how well AI agents perform at machine learning engineeringgithub.com
- Competitive programming with AlphaCode — Google DeepMinddeepmind.com
- As Rocks May Think | Eric Jangevjang.com
- PostTrainBenchposttrainbench.com
- BigO(Bench) -- Can LLMs Generate Code with Controlled Time and Space Complexity?alphaxiv.org