flâneur

Scaling Video Pretraining with Imagination Models — Induction Labs

inductionlabs.com · 1,974 words · saved by 1 readers

We introduce imagination models, a foundation model architecture that unlocks scalable learning from internet video. These models learn to imagine the future in a learned representation space. During pretraining, imagination models implicitly learn to act despite seeing no action labels. Afterwards, reinforcement learning can be scaled to consistently improve their competence.

Internet video contains millions of hours of people using computers, acting in the physical world, interacting with one another, and performing skilled work. We introduce imagination models, a simple foundation model architecture that unlocks scalable learning from this internet video. These models learn to imagine the future in a learned representation space. During pretraining, imagination models implicitly learn to act despite seeing no action labels. Afterwards, reinforcement learning can be scaled to consistently improve their competence. We test the imagination model architecture…

saved by

related reading