flâneur — a map of the web's best reading

New report: Evaluating Language-Model Agents on Realistic Autonomous Tasks - METR

metr.org · 1,690 words · saved by 1 readers

We have just released our first public report. It introduces methodology for assessing the capacity of LLM agents to acquire resources, create copies of themselves, and adapt to novel challenges they encounter in the wild. ARC Evals develops methods for evaluating the safety of large language models (LLMs) in order to provide early warnings of models with dangerous capabilities. We have public partnerships with Anthropic and OpenAI to evaluate their AI systems, and are exploring other partnerships as well. We have just released our first public report on these evaluations. It introduces methodology for assessing the capacity of LLM agents to acquire resources, create copies of themselves, and adapt to novel challenges they encounter in the wild. We refer to these capacities as “autonomous replication and adaptation,” or ARA. We see this as a pilot study of the sort of evaluations that will be necessary to ensure the safe development and deployment of LLMs larger than those that have be

New report: Evaluating Language-Model Agents on Realistic Autonomous Tasks - METR Our Work Research Notes Updates Risk Assessment About Donate Careers Search --> Our Work Research Notes Updates Risk Assessment About Donate Careers Menu × New report: Evaluating Language-Model Agents on Realistic Autonomous Tasks CONTRIBUTORS Beth Barnes DATE July 31, 2023 SHARE Copy Link Citation BibTeX Citation × @misc { metr-2023-new-report , title = {New report: Evaluating Language-Model Agents on Realistic Autonomous Tasks} , author = {Beth Barnes} , howpublished = {\url{https://metr.org/blog/20

Explore this link on the map →

related reading