PostTrainBench: Measuring AI Ability to Perform LLM Post-Training
PostTrainBench is our new benchmark that measures whether AI agents can successfully post-train language models—a key capability for AI R&D automation. Each agent is based on a frontier LLM and a native CLI agent scaffold. For example, we use Claude Code for Opus 4.5 and Codex CLI for GPT-5.2. We give each agent small base LLM, an H100 GPU, and 10 hours to improve model performance through fine-tuning. GPT-5.1 Codex Max achieves the best results: 34.9% average benchmark performance vs. 20.1% for the second-best LLM (Claude Opus 4.5). However, a significant gap remains when compared to human post-trained LLMs (61.8%). Code and leaderboard available at PostTrainBench.com and GitHub. As LLM agents become more capable, a natural question emerges: can they perform AI research and development autonomously? One concrete way to measure this is to test whether agents can successfully post-train (fine-tune) language models. Thanks for reading AI Safety and Alignment Group! Subscribe for free to
PostTrainBench: Measuring AI Ability to Perform LLM Post-Training We expect this to be an important indicator for AI R&D automation as it unfolds over the next few years Maksym Andriushchenko , Ben Rank , and Hardik Bhatnagar Jan 06, 2026 7 1 Share TLDR : PostTrainBench is our new benchmark that measures whether AI agents can successfully post-train language models—a key capability for AI R&D automation. Each agent is based on a frontier LLM and a native CLI agent scaffold . For example, we use Claude Code for Opus 4.5 and Codex CLI for GPT-5.2. We give each agent small base LLM, an H100 GPU,
Explore this link on the map →related reading
- PostTrainBenchposttrainbench.com
- [2603.08640] PostTrainBench: Can LLM Agents Automate LLM Post-Training?arxiv.org
- Composer2.pdfcursor.com
- AI in 2025: gestalt — LessWronglesswrong.com
- gpt-4.pdfcdn.openai.com
- 2025: The year in LLMssimonwillison.net
- AINews | AINewsnews.smol.ai
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- Evaluating frontier AI R&D capabilities of language model agents against human experts - METRmetr.org
- Building Effective AI Agents \ Anthropicanthropic.com
- GitHub - METR/RE-Bench · GitHubgithub.com
- GitHub - karpathy/autoresearch: AI agents running research on single-GPU nanochat training automatically · GitHubgithub.com