flâneur — a map of the web's best reading

Evaluating frontier AI R&D capabilities of language model agents against human experts - METR

metr.org · 4,445 words · saved by 1 readers

We’re releasing RE-Bench, a new benchmark for measuring the performance of humans and frontier model agents on ML research engineering tasks. We also share data from 71 human expert attempts and results for Anthropic’s Claude 3.5 Sonnet and OpenAI’s o1-preview.

Evaluating frontier AI R&D capabilities of language model agents against human experts - METR Our Work Research Notes Updates Risk Assessment About Donate Careers Search --> Our Work Research Notes Updates Risk Assessment About Donate Careers Menu × Evaluating frontier AI R&D capabilities of language model agents against human experts DATE November 22, 2024 SHARE Copy Link Citation BibTeX Citation × @misc { metr-2024-evaluating-r-d-capabilities-of-llms , title = {Evaluating frontier AI R&D capabilities of language model agents against human experts} , author = {METR} , howpublished

Explore this link on the map →

related reading