flâneur

𝙔𝙪𝙢𝙚

stdstu12.github.io · 414 words · saved by 1 readers

Yume aims to use images, text, or videos to create an interactive, realistic, and dynamic world, which allows exploration and control using peripheral devices or neural signals. In this report, we present a preview version of Yume, which creates a dynamic world from an input image and allows exploration of the world using keyboard actions. To achieve this high-fidelity and interactive video world generation, we introduce a well-designed framework, which consists of four main components, including camera motion quantization, video generation architecture, advanced sampler, and model acceleration. First, we quantize camera motions for stable training and user-friendly interaction using keyboard inputs. Then, we introduce the Masked Video Diffusion Transformer~(MVDT) with a memory module for infinite video generation in an autoregressive manner. After that, training-free Anti-Artifact Mechanism (AAM) and Time Travel Sampling based on Stochastic Differential Equations (TTS-SDE) are introdu

Zhen Li1, Chuanhao Li1, Xiaojie Xu1, Kaining Ying2, Tong He1, Jiangmiao Pang1, Yu Qiao1, Kaipeng Zhang1,3†‡ 1Shanghai AI Laboratory, 2Fudan University, 3Shanghai Innovation Institute We are looking for collaboration and self-motivated interns. Contact: kp_zhang@foxmail.com. Abstract Recent approaches have demonstrated the promise of using diffusion models to generate interactive and explorable worlds. However, most of these methods face critical challenges such as excessively large parameter sizes, reliance on lengthy inference steps, and rapidly growing historical context, which…

saved by

related reading