Limits to narrow LLM complementarity @ osmarks' website
osmarks.net · 1,477 words · saved by 1 readers
Taste is probably not the bottleneck.
Since late 2024, frontier LLM systems have been trained with high-compute reinforcement learning on outcome rewards[1], as opposed to the previous dominance of reinforcement learning from human feedback (smaller-scale due to the need for labelled data) and self-supervised pretraining. Broadly, this makes them better at "doing tasks" much more quickly and cheaply[2] than scaling pretraining would, but scale improves everything and this is not so general. This has led to some predictions along the lines of "taste[3] is the bottleneck". I don't believe that this has any more long-term validity…
saved by
related reading
- How can LLM RL Work Despite Information-Theoretic Inefficiencyberen.io
- 2025 LLM Year in Review – karpathykarpathy.bearblog.dev
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- The Extreme Inefficiency of RL for Frontier Models - Toby Ordtobyord.com
- What We’ve Learned From A Year of Building with LLMs – Applied LLMsapplied-llms.org
- AI in 2025: gestalt — LessWronglesswrong.com
- Animals vs Ghosts – karpathykarpathy.bearblog.dev
- GenAI Handbookgenai-handbook.github.io
- Thoughts on AI progress (Dec 2025)substack.com
- The bitter lesson of LLM evalsparsed.com
- Thoughts on AI progress (Dec 2025)dwarkesh.com