openai/mle-bench: MLE-bench is a benchmark for measuring how well AI agents perform at machine learning engineering ·
github.com · 2,315 words · saved by 1 readers
MLE-bench is a benchmark for measuring how well AI agents perform at machine learning engineering
Code for the paper "MLE-Bench: Evaluating Machine Learning Agents on Machine Learning Engineering". We have released the code used to construct the dataset, the evaluation logic, as well as the agents we evaluated for this benchmark. Leaderboard Update (04-24-2026): We are currently not taking any new submissions to the leaderboard while we develop an improved process for ensuring submissions are fair and comparable. We will share updates on this process in the future. Agent LLM(s) used Low == Lite (%) Medium (%) High (%) All (%) Running Time (hours) Date Source Code Available Grading…
saved by
related reading
- PostTrainBenchposttrainbench.com
- Composer2.pdfcursor.com
- Demystifying evals for AI agents \ Anthropicanthropic.com
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- ALE-Bench: A Benchmark for Long-Horizon Objective-Driven Algorithm Engineering | alphaXivalphaxiv.org
- GitHub - METR/RE-Bench · GitHubgithub.com
- Evaluating frontier AI R&D capabilities of language model agents against human experts - METRmetr.org
- TERMINAL-BENCHtbench.ai
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- FrontierSWEfrontierswe.com
- 2025: The year in LLMssimonwillison.net
- Open-Source Agentic Inference Benchmark | InferenceXinferencex.semianalysis.com