A General Goal-Conditioned Minecraft Model - Pantograph
We develop a simple method for learning goal-directed behavior from pretraining on internet-scale video, and use it to train a model capable of achieving diverse and out-of-distribution goals in Minecraft. These goals are first seen at inference time, without training on any of them specifically.
At Pantograph, we're working on training fully general robotics models that can act autonomously for hours at a time. Especially in robotics, it's difficult to get diverse data at scale. Learning to act from internet video data could allow models to scale with compute, rather than being limited by small action datasets. In this work, we develop a simple method for learning goal-directed behavior through pretraining on internet-scale video. Usually, goal-directedness is taught in a post-training phase, which limits the extent to which it can generalize. Here, we learn goal-directedness during p
saved by
related reading
- General Instinct | Any frontier model. Any edge device.general-instinct.com
- RoboTTT: Context Scaling for Robot Policiesresearch.nvidia.com
- [2206.11795] Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online Videosarxiv.org
- Genie: Generative Interactive Environmentsdeepmind.google
- The First Fully General Computer Action Model | blogsi.inc
- How to train a frontier-level world modelnext-state.github.io
- [2509.24527] Training Agents Inside of Scalable World Modelsarxiv.org
- Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models | NVIDIA Technical Blogdeveloper.nvidia.com
- [2306.00937] STEVE-1: A Generative Model for Text-to-Behavior in Minecraftarxiv.org
- Causal Video Models Are Data-Efficient Robot Policy Learners | Rhoda AIrhoda.ai
- Scaling Video Pretraining with Imagination Modelsinductionlabs.com
- World Models | Rohit Bandarurohitbandaru.github.io