flâneur — a map of the web's best reading

Automated Weak-to-Strong Researcher

alignment.anthropic.com · 4,917 words · saved by 15 readers

We built autonomous AI agents that propose ideas, run experiments, and iterate on an open research problem: how to train a strong model using only a weaker model's supervision. These agents outperform human researchers, suggesting that automating this kind of research is already practical. Research partially done as part of the Anthropic Fellows Program. Code available here: https://github.com/safety-research/automated-w2s-research. Sign-offs (tick the box next to your name when you are happy for it to go out): Today’s alignment progress is bottlenecked by human researchers. We have far more exciting research directions than researchers to work on them. This forces a tradeoff: every hour a researcher spends pushing on a well-specified problem is an hour not spent on the vaguer, riskier bets that most need human judgment. If we can hand off the former, we free ourselves for the latter. To address this bottleneck, we build a Claude-powered Automated Alignment Researcher (AAR) that turns

Automated Weak-to-Strong Researcher Alignment Science Blog Automated Weak-to-Strong Researcher Jiaxin Wen*, Liang Qiu*, Joe Benton, Jan Hendrik Kirchner, Jan Leike TL;DR: We built autonomous AI agents that propose ideas, run experiments, and iterate on an open research problem: how to train a strong model using only a weaker model's supervision. These agents outperform human researchers, suggesting that automating this kind of research is already practical. Research partially done as part of the Anthropic Fellows Program. Code available here: https://github.com/safety-research/automated-w2s-re

Explore this link on the map →

saved by

related reading