✳flâneur — a map of the web's best reading
I Let AI Agents Train Their Own Models. Here's What Actually Happened. | Hamza Mostafa
hamzamostafa.com · 1,685 words · saved by 1 readers
Two frontier agents, a pile of bugs, and a reality check on the future of autonomous AI research.
All posts TL;DR: I built a system called Tinkerer that lets frontier AI agents (Claude Code, OpenAI Codex) autonomously fine-tune language models — no human in the loop. After 100+ experiments, the best run produced near-perfect arithmetic from a 3B model. The worst runs burned 10+ hours of compute because neither agent noticed a broken learning rate scheduler. My main takeaway: these agents can execute training pipelines, but they can't yet do ML research. Training is execution. Research is judgment. They're good at the first, still developing the second. The Future Everyone Is Talking About
Explore this link on the map →related reading
- AI 2027ai-2027.com
- Automated Weak-to-Strong Researcheralignment.anthropic.com
- When AI builds itself \ Anthropicanthropic.com
- Composer2.pdfcursor.com
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- If you haven’t recently used Claude Code*, you might not understand where AI is atdavidpreichert.substack.com
- Building Effective AI Agents \ Anthropicanthropic.com
- Why Tool AIs Want to Be Agent AIs · Gwern.netgwern.net
- AI 2027ai-2027.com
- [2603.08640] PostTrainBench: Can LLM Agents Automate LLM Post-Training?arxiv.org
- Effective harnesses for long-running agents \ Anthropicanthropic.com
- LLM Powered Autonomous Agents | Lil'Loglilianweng.github.io