METR/RE-Bench ·
We intend for these tasks to serve as example evaluation material aimed at measuring the autonomous AI R&D capabilities of AI agents. For more information, see the full paper. All the tasks in this repo conform to the METR Task Standard. The METR Task Standard is our attempt at defining a common format for tasks. We hope that this format will help facilitate easier task sharing and agent evaluation. See the setup guide for getting started running this task suite with Vivaria and our open source agent scaffolding. This repo is licensed under the MIT License. See LICENSE for details. In addition to this license we informally request that people take reasonable steps to keep these tasks out of LLM training data and avoid overfitting, including: To help protect solution information from ending up in training data, some tasks have files that are only available via password-protected zips. We would like to ask that people do not publish un-protected solutions to these tasks. If you accidenta
RE-Bench Evaluating frontier AI R&D capabilities of language model agents against human experts We intend for these tasks to serve as example evaluation material aimed at measuring the autonomous AI R&D capabilities of AI agents. For more information, see the full paper . METR Task Standard All the tasks in this repo conform to the METR Task Standard . The METR Task Standard is our attempt at defining a common format for tasks. We hope that this format will help facilitate easier task sharing and agent evaluation. See the setup guide for getting started running this task suite with Vivaria and
Explore this link on the map →saved by
related reading
- Evaluating frontier AI R&D capabilities of language model agents against human experts - METRmetr.org
- Composer2.pdfcursor.com
- gpt-4.pdfcdn.openai.com
- AI in 2025: gestalt — LessWronglesswrong.com
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- PostTrainBenchposttrainbench.com
- Measuring AI Ability to Complete Long Tasks - METRmetr.org
- Patterns for Building LLM-based Systems & Productseugeneyan.com
- GitHub - openai/mle-bench: MLE-bench is a benchmark for measuring how well AI agents perform at machine learning engineering · GitHubgithub.com
- Frontier Risk Report (February to March 2026) - METRmetr.org
- The bitter lesson of LLM evalsparsed.com
- Lapis Labslapis.rocks