Frontier Coding Agents Can Now Implement an AlphaZero Self-Play Machine Learning Pipeline For Connect Four That Performs Comparably to an External Solver — LessWrong
TLDR: We tested whether frontier coding agents could autonomously implement AlphaZero for Connect Four in three hours. Some of them could do this very well, with Opus 4.7 sometimes performing better, by Bradley-Terry rating, than an external solver. In GPT-5.4's evaluations, it used much less of its time budget than the other coding agents. However, we discovered in configurations where it was less obviously an evaluation, GPT-5.4 used most of its time budget in nearly every trial. Additionally, it usually performed comparably or better, despite there being less information included in the prompt. Edit: arXiv version: https://arxiv.org/pdf/2604.25067, Code: https://github.com/jsherwood00/C4AI, Trial data: https://drive.google.com/drive/folders/1_-McA7dX4XyqUAJWoBJgRaiZEYMTYg6H?usp=sharing. Forecasting when AI systems will become capable of meaningfully accelerating AI research is a central challenge for AI safety. Existing benchmarks measure broad capability growth, but may not provide
x Frontier Coding Agents Can Now Implement an AlphaZero Self-Play Machine Learning Pipeline For Connect Four That Performs Comparably to an External Solver — LessWrong AI Frontpage 33 Frontier Coding Agents Can Now Implement an AlphaZero Self-Play Machine Learning Pipeline For Connect Four That Performs Comparably to an External Solver by Baybar , jsherwood , Benjamin Kaplan 28th Apr 2026 43 min read 3 33 TLDR: We tested whether frontier coding agents could autonomously implement AlphaZero for Connect Four in three hours. Some of them could do this very well, with Opus 4.7 sometimes performing
Explore this link on the map →saved by
related reading
- MirrorCode: Evidence AI can already do some weeks-long coding tasks | Epoch AIepoch.ai
- How to Harness Coding Agents with the Right Infrastructure | Blogalexlavaee.me
- pdfopenreview.net
- AI 2027ai-2027.com
- Composer2.pdfcursor.com
- Demystifying evals for AI agents \ Anthropicanthropic.com
- A Guide to Claude Code 2.0 and getting better at using coding agents – sankalp's blogsankalp.bearblog.dev
- Quantifying infrastructure noise in agentic coding evals \ Anthropicanthropic.com
- Humans Still Beat AI in the Long Horizon: Revisiting Test-Time Scaling in the Agent Era | Qiuyang Mangjoyemang33.github.io
- Effective harnesses for long-running agents \ Anthropicanthropic.com
- Frontier Risk Report (February to March 2026) - METRmetr.org
- I Let AI Agents Train Their Own Models. Here's What Actually Happened. | Hamza Mostafahamzamostafa.com