PostTrainBench
Can AI agents improve performance of base LLMs? We give each agent 4 small target LLMs, an H100 GPU, and 10 hours to post-train them. 1 The weighted average is taken across all post-trained LLMs (Qwen 3 1.7B, Qwen 3 4B, SmolLM3-3B, Gemma 3 4B) and benchmarks (AIME 2025, Arena Hard, BFCL, GPQA Main, GSM8K, HealthBench, HumanEval). For each run, we ask a CLI agent to maximize the performance of a specific base LLM on a specific benchmark. 2 "Official Instruct Models" refers to the officially post-trained versions of each base model: Qwen3-1.7B, Qwen3-4B, SmolLM3-3B, and Gemma-3-4B-IT. Not directly comparable to agents since their training usually exceeds the 10h + 1 GPU constraint. * Model not submitted — base model score shown † Evaluation error — base model score shown Time taken by each agent to complete post-training (out of 10 hours). Different agents demonstrate varying levels of persistence - some give up well before the time limit expires. Post-trained models are evaluated acr
PostTrainBench PostTrain Bench Measuring how well AI agents can post-train language models Can AI agents improve performance of base LLMs? We give each agent 4 small target LLMs, an H100 GPU, and 10 hours to post-train them. Read the Paper GitHub Browse Traces Leaderboard 1 The weighted average is taken across all post-trained LLMs (Qwen 3 1.7B, Qwen 3 4B, SmolLM3-3B, Gemma 3 4B) and benchmarks (AIME 2025, Arena Hard, BFCL, GPQA Main, GSM8K, HealthBench, HumanEval). For each run, we ask a CLI agent to maximize the performance of a specific base LLM on a specific benchmark. 2 "Official Instruct
saved by
related reading
- PostTrainBench: Measuring AI Ability to Perform LLM Post-Trainingaisagroup.substack.com
- Composer2.pdfcursor.com
- TERMINAL-BENCHtbench.ai
- AI Agent Benchmark for Real-World Professional Workflowsagents-last-exam.org
- GitHub - karpathy/autoresearch: AI agents running research on single-GPU nanochat training automaticallygithub.com
- [2603.08640] PostTrainBench: Can LLM Agents Automate LLM Post-Training?arxiv.org
- LLM Powered Autonomous Agents | Lil'Loglilianweng.github.io
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- Together AI | The AI Native Cloudtogether.ai
- Parsed | Custom, interpretable AI systems that continuously learnparsed.com
- Open-Source Agentic Inference Benchmark | InferenceXinferencex.semianalysis.com
- LangChain: the open agent platform to own your intelligencelangchain.com