flâneur — a map of the web's best reading

METR/RE-Bench ·

github.com · 831 words · saved by 1 readers

We intend for these tasks to serve as example evaluation material aimed at measuring the autonomous AI R&D capabilities of AI agents. For more information, see the full paper. All the tasks in this repo conform to the METR Task Standard. The METR Task Standard is our attempt at defining a common format for tasks. We hope that this format will help facilitate easier task sharing and agent evaluation. See the setup guide for getting started running this task suite with Vivaria and our open source agent scaffolding. This repo is licensed under the MIT License. See LICENSE for details. In addition to this license we informally request that people take reasonable steps to keep these tasks out of LLM training data and avoid overfitting, including: To help protect solution information from ending up in training data, some tasks have files that are only available via password-protected zips. We would like to ask that people do not publish un-protected solutions to these tasks. If you accidenta

RE-Bench Evaluating frontier AI R&D capabilities of language model agents against human experts We intend for these tasks to serve as example evaluation material aimed at measuring the autonomous AI R&D capabilities of AI agents. For more information, see the full paper . METR Task Standard All the tasks in this repo conform to the METR Task Standard . The METR Task Standard is our attempt at defining a common format for tasks. We hope that this format will help facilitate easier task sharing and agent evaluation. See the setup guide for getting started running this task suite with Vivaria and

Explore this link on the map →

saved by

related reading