Language model harnesses are compositional generalizers | Alex L. Zhang
Harnesses can lead to compositional generalization: we observe a property in training RLMs, in which similarly structured tasks are viewed as isomorphic and all individual LM calls in the harness become in-distribution.
Modern post-training has become a brute-force paradigm of curating ever more environments and ever longer training horizons. In large part, this is because frontier Transformers are still poor at compositional generalization, the ability to solve unseen problems by composing familiar ones. Unless our models compose the individual lessons they learn, scaling will have slower returns than it should, as every new domain will demand its own investment in the form of training data. Training data is not the only lever, of course. For the past few years, we’ve attacked harder tasks by scaffolding…
saved by
related reading
- Alex L. Zhangalexzhang13.github.io
- Getting Caught Up to Modern LLM Research | Samarth Goeldev.samarthgoel.com
- Composer2.pdfcursor.com
- To Understand Language is to Understand Generalization | Eric Jangevjang.com
- Just Ask for Generalization | Eric Jangevjang.com
- GenAI Handbookgenai-handbook.github.io
- A Steerable Model with Emergent Capabilitiespi.website
- The upcoming GPT-3 moment for RL | Mechanize, Inc.mechanize.work
- 2305.18654arxiv.org
- Parameter Golf Research Gardengolf.agustif.com
- [2310.16028] What Algorithms can Transformers Learn? A Study in Length Generalizationarxiv.org
- [2506.19733] Breaking Barriers: Do Reinforcement Post Training Gains Transfer To Unseen Domains?arxiv.org