flâneur

Language model harnesses are compositional generalizers | Alex L. Zhang

alexzhang13.github.io · 5,073 words · saved by 1 readers

Harnesses can lead to compositional generalization: we observe a property in training RLMs, in which similarly structured tasks are viewed as isomorphic and all individual LM calls in the harness become in-distribution.

Modern post-training has become a brute-force paradigm of curating ever more environments and ever longer training horizons. In large part, this is because frontier Transformers are still poor at compositional generalization, the ability to solve unseen problems by composing familiar ones. Unless our models compose the individual lessons they learn, scaling will have slower returns than it should, as every new domain will demand its own investment in the form of training data. Training data is not the only lever, of course. For the past few years, we’ve attacked harder tasks by scaffolding…

saved by

related reading