PostTrainBench
Can AI agents improve performance of base LLMs? We give each agent 4 small target LLMs, an H100 GPU, and 10 hours to post-train them. 1 The weighted average is taken across all post-trained LLMs (Qwen 3 1.7B, Qwen 3 4B, SmolLM3-3B, Gemma 3 4B) and benchmarks (AIME 2025, Arena Hard, BFCL, GPQA Main, GSM8K, HealthBench, HumanEval). For each run, we ask a CLI agent to maximize the performance of a specific base LLM on a specific benchmark. 2 "Official Instruct Models" refers to the officially post-trained versions of each base model: Qwen3-1.7B, Qwen3-4B, SmolLM3-3B, and Gemma-3-4B-IT. Not directly comparable to agents since their training usually exceeds the 10h + 1 GPU constraint. * Model not submitted — base model score shown † Evaluation error — base model score shown Time taken by each agent to complete post-training (out of 10 hours). Different agents demonstrate varying levels of persistence - some give up well before the time limit expires. Post-trained models are evaluated acr
PostTrainBench PostTrain Bench Measuring how well AI agents can post-train language models Can AI agents improve performance of base LLMs? We give each agent 4 small target LLMs, an H100 GPU, and 10 hours to post-train them. Read the Paper GitHub Browse Traces Leaderboard 1 The weighted average is taken across all post-trained LLMs (Qwen 3 1.7B, Qwen 3 4B, SmolLM3-3B, Gemma 3 4B) and benchmarks (AIME 2025, Arena Hard, BFCL, GPQA Main, GSM8K, HealthBench, HumanEval). For each run, we ask a CLI agent to maximize the performance of a specific base LLM on a specific benchmark. 2 "Official Instruct
Explore this link on the map →saved by
related reading
- AI Agent Benchmark for Real-World Professional Workflowsagents-last-exam.org
- Composer2.pdfcursor.com
- AI Benchmark Leaderboards & Model Evals | BenchmarkListbenchmarklist.com
- PostTrainBench: Measuring AI Ability to Perform LLM Post-Trainingaisagroup.substack.com
- The bitter lesson of LLM evalsparsed.com
- [2603.08640] PostTrainBench: Can LLM Agents Automate LLM Post-Training?arxiv.org
- LLM Powered Autonomous Agents | Lil'Loglilianweng.github.io
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- Patterns for Building LLM-based Systems & Productseugeneyan.com
- GitHub - karpathy/autoresearch: AI agents running research on single-GPU nanochat training automatically · GitHubgithub.com
- GitHub - openai/mle-bench: MLE-bench is a benchmark for measuring how well AI agents perform at machine learning engineering · GitHubgithub.com
- Agent Evaluation: A Detailed Guidecameronrwolfe.substack.com