Automated Weak-to-Strong Researcher
We built autonomous AI agents that propose ideas, run experiments, and iterate on an open research problem: how to train a strong model using only a weaker model's supervision. These agents outperform human researchers, suggesting that automating this kind of research is already practical. Research partially done as part of the Anthropic Fellows Program. Code available here: https://github.com/safety-research/automated-w2s-research. Sign-offs (tick the box next to your name when you are happy for it to go out): Today’s alignment progress is bottlenecked by human researchers. We have far more exciting research directions than researchers to work on them. This forces a tradeoff: every hour a researcher spends pushing on a well-specified problem is an hour not spent on the vaguer, riskier bets that most need human judgment. If we can hand off the former, we free ourselves for the latter. To address this bottleneck, we build a Claude-powered Automated Alignment Researcher (AAR) that turns
Automated Weak-to-Strong Researcher Alignment Science Blog Automated Weak-to-Strong Researcher Jiaxin Wen*, Liang Qiu*, Joe Benton, Jan Hendrik Kirchner, Jan Leike TL;DR: We built autonomous AI agents that propose ideas, run experiments, and iterate on an open research problem: how to train a strong model using only a weaker model's supervision. These agents outperform human researchers, suggesting that automating this kind of research is already practical. Research partially done as part of the Anthropic Fellows Program. Code available here: https://github.com/safety-research/automated-w2s-re
Explore this link on the map →saved by
- Karthik Suresh
- Yixiong Hao
- Asher P
- Uzay Girit
- Lydia Nottingham
- Dhruv Sheth
- Abhay Sheshadri
- Kushal Thaman
- Jo J.
- Ishan Mukherjee
- Jeremy Kintana
- Will Anderson
related reading
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- AI 2027ai-2027.com
- [2603.08640] PostTrainBench: Can LLM Agents Automate LLM Post-Training?arxiv.org
- [2604.22082] Removing Sandbagging in LLMs by Training with Weak Supervisionarxiv.org
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Tips for Empirical Alignment Research — AI Alignment Forumalignmentforum.org
- I am worried about near-term non-LLM AI developments — LessWronglesswrong.com
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- Autoresearch Is Not About Training Models. It Is About What Happens When Agents Get a Scoreboard | by Krish | AI Sutraaisutra.com
- Danger, AI Scientist, Danger - by Zvi Mowshowitzthezvi.substack.com
- Import AI 455: AI systems are about to start building themselves.importai.substack.com