Evaluating frontier AI R&D capabilities of language model agents against human experts - METR
We’re releasing RE-Bench, a new benchmark for measuring the performance of humans and frontier model agents on ML research engineering tasks. We also share data from 71 human expert attempts and results for Anthropic’s Claude 3.5 Sonnet and OpenAI’s o1-preview.
Evaluating frontier AI R&D capabilities of language model agents against human experts - METR Our Work Research Notes Updates Risk Assessment About Donate Careers Search --> Our Work Research Notes Updates Risk Assessment About Donate Careers Menu × Evaluating frontier AI R&D capabilities of language model agents against human experts DATE November 22, 2024 SHARE Copy Link Citation BibTeX Citation × @misc { metr-2024-evaluating-r-d-capabilities-of-llms , title = {Evaluating frontier AI R&D capabilities of language model agents against human experts} , author = {METR} , howpublished
Explore this link on the map →related reading
- GitHub - METR/RE-Bench · GitHubgithub.com
- Composer2.pdfcursor.com
- Automated Weak-to-Strong Researcheralignment.anthropic.com
- AI in 2025: gestalt — LessWronglesswrong.com
- Measuring AI Ability to Complete Long Tasks - METRmetr.org
- PostTrainBenchposttrainbench.com
- Demystifying evals for AI agents \ Anthropicanthropic.com
- GitHub - openai/mle-bench: MLE-bench is a benchmark for measuring how well AI agents perform at machine learning engineering · GitHubgithub.com
- An Apple-Picking Model of AI R&D | Tom Cunningham – Tom Cunninghamtecunningham.github.io
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- Agent Evaluation: A Detailed Guidecameronrwolfe.substack.com
- [2603.08640] PostTrainBench: Can LLM Agents Automate LLM Post-Training?arxiv.org