flâneur — a map of the web's best reading

PostTrainBench

posttrainbench.com · 1,232 words · saved by 2 readers

Can AI agents improve performance of base LLMs? We give each agent 4 small target LLMs, an H100 GPU, and 10 hours to post-train them. 1 The weighted average is taken across all post-trained LLMs (Qwen 3 1.7B, Qwen 3 4B, SmolLM3-3B, Gemma 3 4B) and benchmarks (AIME 2025, Arena Hard, BFCL, GPQA Main, GSM8K, HealthBench, HumanEval). For each run, we ask a CLI agent to maximize the performance of a specific base LLM on a specific benchmark. 2 "Official Instruct Models" refers to the officially post-trained versions of each base model: Qwen3-1.7B, Qwen3-4B, SmolLM3-3B, and Gemma-3-4B-IT. Not directly comparable to agents since their training usually exceeds the 10h + 1 GPU constraint. * Model not submitted — base model score shown    † Evaluation error — base model score shown Time taken by each agent to complete post-training (out of 10 hours). Different agents demonstrate varying levels of persistence - some give up well before the time limit expires. Post-trained models are evaluated acr

PostTrainBench PostTrain Bench Measuring how well AI agents can post-train language models Can AI agents improve performance of base LLMs? We give each agent 4 small target LLMs, an H100 GPU, and 10 hours to post-train them. Read the Paper GitHub Browse Traces Leaderboard 1 The weighted average is taken across all post-trained LLMs (Qwen 3 1.7B, Qwen 3 4B, SmolLM3-3B, Gemma 3 4B) and benchmarks (AIME 2025, Arena Hard, BFCL, GPQA Main, GSM8K, HealthBench, HumanEval). For each run, we ask a CLI agent to maximize the performance of a specific base LLM on a specific benchmark. 2 "Official Instruct

Explore this link on the map →

saved by

related reading