Scaling Video Pretraining with Imagination Models — Induction Labs
We introduce imagination models, a foundation model architecture that unlocks scalable learning from internet video. These models learn to imagine the future in a learned representation space. During pretraining, imagination models implicitly learn to act despite seeing no action labels. Afterwards, reinforcement learning can be scaled to consistently improve their competence.
Internet video contains millions of hours of people using computers, acting in the physical world, interacting with one another, and performing skilled work. We introduce imagination models, a simple foundation model architecture that unlocks scalable learning from this internet video. These models learn to imagine the future in a learned representation space. During pretraining, imagination models implicitly learn to act despite seeing no action labels. Afterwards, reinforcement learning can be scaled to consistently improve their competence. We test the imagination model architecture…
saved by
related reading
- The First Fully General Computer Action Model | blogsi.inc
- [2206.11795] Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online Videosarxiv.org
- Explore | alphaXivalphaxiv.org
- Generalist - GEN-0 / Embodied Foundation Models That Scale with Physical Interactiongeneralistai.com
- [2509.24527] Training Agents Inside of Scalable World Modelsarxiv.org
- World Models: Computing the Uncomputablenotboring.co
- 45d74e190008c7bff2845ffc8e3facd3-Paper-Conference.pdfproceedings.iclr.cc
- The Model That Dreams the Worldmoe-capital.com
- Learning to Act without Actionsarxiv.org
- Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models | NVIDIA Technical Blogdeveloper.nvidia.com
- General Instinct | Any frontier model. Any edge device.general-instinct.com
- World Models | Rohit Bandarurohitbandaru.github.io