flâneur

openai/mle-bench: MLE-bench is a benchmark for measuring how well AI agents perform at machine learning engineering ·

github.com · 2,315 words · saved by 1 readers

MLE-bench is a benchmark for measuring how well AI agents perform at machine learning engineering

Code for the paper "MLE-Bench: Evaluating Machine Learning Agents on Machine Learning Engineering". We have released the code used to construct the dataset, the evaluation logic, as well as the agents we evaluated for this benchmark. Leaderboard Update (04-24-2026): We are currently not taking any new submissions to the leaderboard while we develop an improved process for ensuring submissions are fair and comparable. We will share updates on this process in the future. Agent LLM(s) used Low == Lite (%) Medium (%) High (%) All (%) Running Time (hours) Date Source Code Available Grading…

saved by

related reading