AI agents can't yet do open-ended AI research
normaltech.ai · 1,365 words · saved by 1 readers
Early evidence from two case studies
The goal of leading AI labs is recursive self-improvement (RSI): the automation of AI research using AI agents. RSI also underpins forecasts of explosive AI progress. How can we assess if we are close to this milestone? One way is to use benchmarks that test if agents can conduct AI research. Given the AI community’s focus on benchmarks, they have been the dominant way to evaluate progress towards RSI. Over the last year, many such evaluations have found that agents are now able to make progress on tasks where success is easily verifiable, prompting speculation that we are on the verge of…
related reading
- AI 2027ai-2027.com
- An Apple-Picking Model of AI R&D | Tom Cunningham – Tom Cunninghamtecunningham.github.io
- Automated Weak-to-Strong Researcheralignment.anthropic.com
- AI 2027ai-2027.com
- Will We See AI with Recursive Self Improvement in 2028? Likely Not.substack.com
- Demystifying evals for AI agents \ Anthropicanthropic.com
- Dwarkesh Podcast | Substackdwarkeshpatel.com
- Thoughts — Jason Weijasonwei.net
- pdfopenreview.net
- Import AI 455: AI systems are about to start building themselves.importai.substack.com
- Evidence on AI R&D Progress from NanoGPT - METRmetr.org
- Ryan Greenblatt – What happens once AI can automate AI research?dwarkesh.com