Language Model Pilot Report - METR
In this report, we explore the ability of language model agents to acquire resources, create copies of themselves, and adapt to novel challenges they encounter in the wild. We refer to this cluster of capabilities as “autonomous replication and adaptation” or ARA. We believe that systems capable of ARA could have wide-reaching and hard-to-anticipate consequences, and that measuring and forecasting ARA may be useful for informing measures around security, monitoring, and alignment. Additionally, once a system is capable of ARA, placing bounds on a system’s capabilities may become significantly more difficult. We construct four simple example agents that combine language models with tools that allow them to take actions in the world. We then evaluate these agents on 12 tasks relevant to ARA. We find that these language model agents can only complete the easiest tasks from this list, although they make some progress on the more challenging tasks. Unfortunately, these evaluations are not a
Language Model Pilot Report - METR Our Work Research Notes Updates Risk Assessment About Donate Careers Search --> Our Work Research Notes Updates Risk Assessment About Donate Careers Menu × Language Model Pilot Report CONTRIBUTORS Megan Kinniment , Lucas Jun Koba Sato , Haoxing Du , Brian Goodrich , Max Hasin , Lawrence Chan , Luke Harold Miles , Tao R Lin , Hjalmar Wijk , Joel Burget , Aaron Ho , Elizabeth Barnes , and Paul Christiano DATE July 31, 2023 SHARE Copy Link Citation BibTeX Citation × @misc { metr-2023-language-model-pilot-report , title = {Language Model Pilot Report}
Explore this link on the map →related reading
- New report: Evaluating Language-Model Agents on Realistic Autonomous Tasks - METRmetr.org
- LLM Powered Autonomous Agents | Lil'Loglilianweng.github.io
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- Automated Weak-to-Strong Researcheralignment.anthropic.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- How fast is AI improving? - AI Digesttheaidigest.org
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- Externalized reasoning oversight: a research direction for language model alignment — AI Alignment Forumalignmentforum.org
- Language Models can Solve Computer Tasksarxiv.org
- Frontier Risk Report (February to March 2026) - METRmetr.org
- Self-Adapting Language Modelsarxiv.org
- pdfopenreview.net