Can Video World Models Track Unobserved World States?
joonghyuk.com · 1,509 words · saved by 1 readers
An empirical study of hidden-state tracking in action-conditioned video world models, using a five-cup Shell Game as a visual analog of S5 state tracking.
Current video models render the visible world remarkably well, yet fail to track its unobserved state. Ground Truth Your browser does not support the video tag. Causal DiT (SWA1) · PSNR 40.4 / Acc 0.12 Causal DiT (SWA1) + LaCT · PSNR 41.1 / Acc 1.00 All models are trained with diffusion forcing on 5-swap episodes. At test time each model receives the clean frames of the opening reveal, then generates all 10 swaps and the final lift from the action sequence alone. Neither wider attention nor more denoising steps help. Proper state tracking requires a carried state and an expressive…
saved by
related reading
- World Action Model Atlasjoeclinton.me
- [2602.10098] VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Modelarxiv.org
- Seoul World Model: Grounding World Simulation Models in a Real-World Metropolisseoul-world-model.github.io
- Video models are zero-shot learners and reasonersarxiv.org
- Video Generation Models Explosion 2024yenchenlin.me
- Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models | NVIDIA Technical Blogdeveloper.nvidia.com
- World Models | Rohit Bandarurohitbandaru.github.io
- State of Vision-Language-Action (VLA) Research at ICLR 2026 – Moritz Reussmbreuss.github.io
- The First Fully General Computer Action Model | blogsi.inc
- World Models: Computing the Uncomputablenotboring.co
- How not to do research - Rajan Agarwalrajan.sh
- [2509.24527] Training Agents Inside of Scalable World Modelsarxiv.org