Automated Weak-to-Strong Researcher
We built autonomous AI agents that propose ideas, run experiments, and iterate on an open research problem: how to train a strong model using only a weaker model's supervision. These agents outperform human researchers, suggesting that automating this kind of research is already practical. Research partially done as part of the Anthropic Fellows Program. Code available here: https://github.com/safety-research/automated-w2s-research. Sign-offs (tick the box next to your name when you are happy for it to go out): Today’s alignment progress is bottlenecked by human researchers. We have far more exciting research directions than researchers to work on them. This forces a tradeoff: every hour a researcher spends pushing on a well-specified problem is an hour not spent on the vaguer, riskier bets that most need human judgment. If we can hand off the former, we free ourselves for the latter. To address this bottleneck, we build a Claude-powered Automated Alignment Researcher (AAR) that turns
Automated Weak-to-Strong Researcher Alignment Science Blog Automated Weak-to-Strong Researcher Jiaxin Wen*, Liang Qiu*, Joe Benton, Jan Hendrik Kirchner, Jan Leike TL;DR: We built autonomous AI agents that propose ideas, run experiments, and iterate on an open research problem: how to train a strong model using only a weaker model's supervision. These agents outperform human researchers, suggesting that automating this kind of research is already practical. Research partially done as part of the Anthropic Fellows Program. Code available here: https://github.com/safety-research/automated-w2s-re
saved by
- Karthik Suresh
- Yixiong Hao
- Asher P
- Uzay Girit
- Lydia Nottingham
- Dhruv Sheth
- Abhay Sheshadri
- Kushal Thaman
- Jo J.
- Ishan Mukherjee
- Jeremy Kintana
- Will Anderson
related reading
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- AI 2027ai-2027.com
- Trending Papers - Hugging Facepaperswithcode.com
- GitHub - karpathy/autoresearch: AI agents running research on single-GPU nanochat training automaticallygithub.com
- [2603.08640] PostTrainBench: Can LLM Agents Automate LLM Post-Training?arxiv.org
- [2604.22082] Removing Sandbagging in LLMs by Training with Weak Supervisionarxiv.org
- As Rocks May Think | Eric Jangevjang.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- The Universe from an Intentional Stancecasparoesterheld.com
- I am worried about near-term non-LLM AI developments — LessWronglesswrong.com
- Predicting Empirical AI Research Outcomes with Language Modelsarxiv.org
- Discovering 108 tricks to accelerate grokkingkindxiaoming.github.io