I Let AI Agents Train Their Own Models. Here's What Actually Happened. | Hamza Mostafa
hamzamostafa.com · 1,685 words · saved by 1 readers
Two frontier agents, a pile of bugs, and a reality check on the future of autonomous AI research.
All posts TL;DR: I built a system called Tinkerer that lets frontier AI agents (Claude Code, OpenAI Codex) autonomously fine-tune language models — no human in the loop. After 100+ experiments, the best run produced near-perfect arithmetic from a 3B model. The worst runs burned 10+ hours of compute because neither agent noticed a broken learning rate scheduler. My main takeaway: these agents can execute training pipelines, but they can't yet do ML research. Training is execution. Research is judgment. They're good at the first, still developing the second. The Future Everyone Is Talking About
related reading
- AI 2027ai-2027.com
- When AI builds itself \ Anthropicanthropic.com
- Automated Weak-to-Strong Researcheralignment.anthropic.com
- GitHub - karpathy/autoresearch: AI agents running research on single-GPU nanochat training automaticallygithub.com
- Composer2.pdfcursor.com
- If you haven’t recently used Claude Code*, you might not understand where AI is atdavidpreichert.substack.com
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- Building Effective AI Agents \ Anthropicanthropic.com
- Trending Papers - Hugging Facepaperswithcode.com
- Effective harnesses for long-running agents \ Anthropicanthropic.com
- [2603.08640] PostTrainBench: Can LLM Agents Automate LLM Post-Training?arxiv.org
- AI 2027ai-2027.com