动察 Beating AI Flash: Bloomberg, citing sources familiar with the matter, reports that ByteDance is developing a real-time spatial video model, with a potential release as early as October, personally overseen by founder Zhang Yiming. Built on Seedance, the model aims to generate content while simultaneously responding to users' voices and actions, creating an interactive virtual world in real time.
The model is planned to be integrated with Pico headsets. According to insiders, it can generate content in the cloud at approximately 20 frames per second with a latency of about 0.05 seconds. By offloading the primary computing to the cloud, the headset's own processing requirements can be reduced. Beyond VR, ByteDance also intends to apply this technology to live streaming, short dramas, and gaming.
This is not a sudden pivot by ByteDance toward world models. As early as June, 36Kr disclosed that world models are one of ByteDance's four key AI priorities for 2026, with a goal of releasing at least one version by year-end, benchmarking against Google's Genie 3. Genie 3 can already generate interactive 720p worlds at 20–24 frames per second and sustain them for several minutes.

