Evaluating frontier AI R&D capabilities of language model agents against human experts - METR
We’re releasing RE-Bench, a new benchmark for measuring the performance of humans and frontier model agents on ML research engineering tasks. We also share data from 71 human expert attempts and results for Anthropic’s Claude 3.5 Sonnet and OpenAI’s o1-preview.
Evaluating frontier AI R&D capabilities of language model agents against human experts - METR Our Work Research Notes Updates Risk Assessment About Donate Careers Search --> Our Work Research Notes Updates Risk Assessment About Donate Careers Menu × Evaluating frontier AI R&D capabilities of language model agents against human experts DATE November 22, 2024 SHARE Copy Link Citation BibTeX Citation × @misc { metr-2024-evaluating-r-d-capabilities-of-llms , title = {Evaluating frontier AI R&D capabilities of language model agents against human experts} , author = {METR} , howpublished
related reading
- GitHub - METR/RE-Bench · GitHubgithub.com
- Composer2.pdfcursor.com
- PostTrainBenchposttrainbench.com
- FrontierSWEfrontierswe.com
- AI in 2025: gestalt — LessWronglesswrong.com
- Measuring AI Ability to Complete Long Tasks - METRmetr.org
- GitHub - openai/mle-bench: MLE-bench is a benchmark for measuring how well AI agents perform at machine learning engineeringgithub.com
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- AINews | AINewsnews.smol.ai
- Frontier Risk Report (February to March 2026) - METRmetr.org
- Are AI benchmarks doomed? - by Anson Ho and Greg Burnhamepochai.substack.com
- Quantifying infrastructure noise in agentic coding evals \ Anthropicanthropic.com