Coding vs thinking — Paradigm 3
We’re interested in the prospects for (presumably safer) narrow AI staying competitive, instead of general systems; Cursor’s Composer coding finetune of Kimi is probably the most intense attempt to specialise a general model: probably more than 10^25 FLOPs of post-training; We find that, compared to its base model, Composer shows…
TL;DR # We’re interested in the prospects for (presumably safer) narrow AI staying competitive, instead of general systems. Cursor’s Composer coding finetune of Kimi is probably the most intense attempt to specialise a general model: probably more than 10^25 FLOPs of post-training. We find that, compared to its base model, Composer shows significant gains (+20% to 60%) on visual reasoning benchmarks, RPG-style games, and (as you’d hope) agentic coding. But, surprisingly, we also saw severe losses (-30% to -40%) on mathematical and scientific reasoning benchmarks. This cuts against the old idea
Explore this link on the map →saved by
related reading
- LLM Chess: Benchmarking Reasoning and Instruction-Following in LLMs through Chessarxiv.org
- Composer2.pdfcursor.com
- AI in 2025: gestalt — LessWronglesswrong.com
- [AINews] Sam Altman's AI Combinator - Latent.Spacelatent.space
- 2025: The year in LLMssimonwillison.net
- AINews | AINewsnews.smol.ai
- My picture of the present in AI — LessWronglesswrong.com
- Inkling: Our Open-Weights Model - Thinking Machines Labthinkingmachines.ai
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- AI progress is about to speed up | Epoch AIepoch.ai
- Coding Models Are Doing Too Much | whnrehiew.github.io
- the-illusion-of-thinking.pdfml-site.cdn-apple.com