flâneur — a map of the web's best reading

Are Video Generation Models World Simulators? · Artificial Cognition

artificialcognition.net · 10,905 words · saved by 1 readers

On February 15th, 2024, OpenAI unveiled Sora – an impressive new deep learning model that can generate videos and images from text prompts. Sora can generate videos up to a minute long, at various resolutions and aspect ratios. While the model is not currently available for testing, cherry-picked results from OpenAI suggest it vastly improves on the previous state of the art.1 OpenAI claims – somewhat pompously – that Sora is a “world simulator.”2 What’s a world simulator, you ask? Good question! Here is OpenAI’s stated motivation for training Sora: We’re teaching AI to understand and simulate the physical world in motion, with the goal of training models that help people solve problems that require real-world interaction.3 The technical report elaborates on OpenAI’s understanding of the theoretical significance of Sora: Our results suggest that scaling video generation models is a promising path towards building general purpose simulators of the physical world.4 These bold statements

This is a deep dive into the claim that OpenAI's new video generation model Sora is a "world simulator." I review what Sora can do, how it works, and what it would mean for it to simulate properties of 3D scenes. The discussion covers the literature on intuitive physics in cognitive science, the polysemous notion of "world model" in machine learning, and interpretability research on image generation models. The upshot is that Sora does not run simulations in the traditional sense, although it may represent physical properties of visual scenes in more limited sense; but behavioral evidence is i

Explore this link on the map →

related reading