flâneur

Limits to narrow LLM complementarity @ osmarks' website

osmarks.net · 1,477 words · saved by 1 readers

Taste is probably not the bottleneck.

Since late 2024, frontier LLM systems have been trained with high-compute reinforcement learning on outcome rewards[1], as opposed to the previous dominance of reinforcement learning from human feedback (smaller-scale due to the need for labelled data) and self-supervised pretraining. Broadly, this makes them better at "doing tasks" much more quickly and cheaply[2] than scaling pretraining would, but scale improves everything and this is not so general. This has led to some predictions along the lines of "taste[3] is the bottleneck". I don't believe that this has any more long-term validity…

saved by

related reading