True Agents Model the World
World modeling during RL trains models to predict their environment in response to their own actions, improving in-domain generalization and efficiency in early ECHO experiments.
True Agents Model the World Pretraining creates simulators , LLMs that can model their environment with high precision. RL only improves the model’s own generations. But true agents should do both at the same time: we don’t want a pure simulator, but neither do we want the LLM to only act, blind to how it will affect its environment. Instead, we want it to learn to predict its environment in response to its own actions, allowing it to more intelligently plan and respond to unexpected world dynamics. Synthetic data during pre- and midtraining achieves similar objectives, but only world modeling
saved by
related reading
- World Models: Computing the Uncomputablenotboring.co
- Explore | alphaXivalphaxiv.org
- World Models | Rohit Bandarurohitbandaru.github.io
- Composer2.pdfcursor.com
- DeepSeek-R1arxiv.org
- LLMs and World Models, Part 1 - by Melanie Mitchellaiguide.substack.com
- As Rocks May Think | Eric Jangevjang.com
- Large Language Model: world models or surface statistics?thegradient.pub
- Language Models, World Models, and Human Model-Buildinglingo.csail.mit.edu
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- How to train a frontier-level world modelnext-state.github.io
- Efficient World Models with Context-Aware Tokenizationarxiv.org