flâneur

Can Video World Models Track Unobserved World States?

joonghyuk.com · 1,509 words · saved by 1 readers

An empirical study of hidden-state tracking in action-conditioned video world models, using a five-cup Shell Game as a visual analog of S5 state tracking.

Current video models render the visible world remarkably well, yet fail to track its unobserved state. Ground Truth Your browser does not support the video tag. Causal DiT (SWA1) · PSNR 40.4 / Acc 0.12 Causal DiT (SWA1) + LaCT · PSNR 41.1 / Acc 1.00 All models are trained with diffusion forcing on 5-swap episodes. At test time each model receives the clean frames of the opening reveal, then generates all 10 swaps and the final lift from the action sequence alone. Neither wider attention nor more denoising steps help. Proper state tracking requires a carried state and an expressive…

saved by

related reading