Frontier Coding Agents Can Now Implement an AlphaZero Self-Play Machine Learning Pipeline For Connect Four That Performs Comparably to an External Solver — LessWrong
TLDR: We tested whether frontier coding agents could autonomously implement AlphaZero for Connect Four in three hours. Some of them could do this very well, with Opus 4.7 sometimes performing better, by Bradley-Terry rating, than an external solver. In GPT-5.4's evaluations, it used much less of its time budget than the other coding agents. However, we discovered in configurations where it was less obviously an evaluation, GPT-5.4 used most of its time budget in nearly every trial. Additionally, it usually performed comparably or better, despite there being less information included in the prompt. Edit: arXiv version: https://arxiv.org/pdf/2604.25067, Code: https://github.com/jsherwood00/C4AI, Trial data: https://drive.google.com/drive/folders/1_-McA7dX4XyqUAJWoBJgRaiZEYMTYg6H?usp=sharing. Forecasting when AI systems will become capable of meaningfully accelerating AI research is a central challenge for AI safety. Existing benchmarks measure broad capability growth, but may not provide
x Frontier Coding Agents Can Now Implement an AlphaZero Self-Play Machine Learning Pipeline For Connect Four That Performs Comparably to an External Solver — LessWrong AI Frontpage 33 Frontier Coding Agents Can Now Implement an AlphaZero Self-Play Machine Learning Pipeline For Connect Four That Performs Comparably to an External Solver by Baybar , jsherwood , Benjamin Kaplan 28th Apr 2026 43 min read 3 33 TLDR: We tested whether frontier coding agents could autonomously implement AlphaZero for Connect Four in three hours. Some of them could do this very well, with Opus 4.7 sometimes performing
saved by
related reading
- MirrorCode: Evidence AI can already do some weeks-long coding tasks | Epoch AIepoch.ai
- How to Harness Coding Agents with the Right Infrastructure | Blogalexlavaee.me
- GitHub - shareAI-lab/learn-claude-code: Bash is all you need - A nano claude code–like 「agent harness」, built from 0 to 1github.com
- AI 2027ai-2027.com
- Composer2.pdfcursor.com
- Frontier Risk Report (February to March 2026) - METRmetr.org
- FrontierSWEfrontierswe.com
- As Rocks May Think | Eric Jangevjang.com
- Quantifying infrastructure noise in agentic coding evals \ Anthropicanthropic.com
- Demystifying evals for AI agents \ Anthropicanthropic.com
- Agent Leaderboards · Which tools coding agents choose · Armaturearmature.tech
- Effective harnesses for long-running agents \ Anthropicanthropic.com