Coding vs thinking — Paradigm 3
We’re interested in the prospects for (presumably safer) narrow AI staying competitive, instead of general systems; Cursor’s Composer coding finetune of Kimi is probably the most intense attempt to specialise a general model: probably more than 10^25 FLOPs of post-training; We find that, compared to its base model, Composer shows…
TL;DR # We’re interested in the prospects for (presumably safer) narrow AI staying competitive, instead of general systems. Cursor’s Composer coding finetune of Kimi is probably the most intense attempt to specialise a general model: probably more than 10^25 FLOPs of post-training. We find that, compared to its base model, Composer shows significant gains (+20% to 60%) on visual reasoning benchmarks, RPG-style games, and (as you’d hope) agentic coding. But, surprisingly, we also saw severe losses (-30% to -40%) on mathematical and scientific reasoning benchmarks. This cuts against the old idea
saved by
related reading
- Composer2.pdfcursor.com
- Training Composer for longer horizons · Cursorcursor.com
- AI in 2025: gestalt — LessWronglesswrong.com
- Introducing Claude Opus 5 \ Anthropicanthropic.com
- As Rocks May Think | Eric Jangevjang.com
- My picture of the present in AI — LessWronglesswrong.com
- Alex L. Zhangalexzhang13.github.io
- Improving Composer through real-time RL · Cursorcursor.com
- Kevin-32B: Multi-Turn RL for Writing CUDA Kernels | Cognitioncognition.ai
- FrontierSWEfrontierswe.com
- [AINews] Sam Altman's AI Combinator - Latent.Spacelatent.space
- AINews | AINewsnews.smol.ai