flâneur

AI agents can't yet do open-ended AI research

normaltech.ai · 1,365 words · saved by 1 readers

Early evidence from two case studies

The goal of leading AI labs is recursive self-improvement (RSI): the automation of AI research using AI agents. RSI also underpins forecasts of explosive AI progress. How can we assess if we are close to this milestone? One way is to use benchmarks that test if agents can conduct AI research. Given the AI community’s focus on benchmarks, they have been the dominant way to evaluate progress towards RSI. Over the last year, many such evaluations have found that agents are now able to make progress on tasks where success is easily verifiable, prompting speculation that we are on the verge of…

related reading