Vivix-W1
Unveiling Vivix-W1: Toward a Streaming-Native Multimodal Interactive Model for Narrative Intelligence
In text-based interaction, humans typically process only ~3–5 tokens per second. In contrast, video-based interaction delivers a continuous multimodal stream composed of tens of thousands of visual tokens, along with synchronized audio tokens, every second. This represents an approximately 10³–10⁴-fold increase in effective token throughput. As a result, real-time video systems must handle enormous information throughput while maintaining millisecond-level latency, continuously perceiving, generating, updating memory, and responding to user input within a unified loop.
Under this regime, the traditional architecture of video generation systems—centered around offline pipelines that follow a static “input prompt → inference → output clip” paradigm—is not well suited to continuous real-time interaction. Real-time interactive video demands a new class of models that are inherently streaming-native, capable of continuously interpreting evolving user intent while simultaneously generating coherent audiovisual content in an ongoing, closed-loop process.
Today, we introduce Vivix-W1, a streaming-native multimodal interactive model with multi-shot storytelling, audio-visual co-generation, and real-time interaction. It is capable of preserving long-horizon consistency across evolving scenes while maintaining fine-grained responsiveness at second-level latency.
Streaming Audio-Visual Omni-Model
Vivix-W1 departs from the standard clip-level diffusion rendering paradigm and natively combines the understanding of interaction signals with a streaming generation architecture.

Vivix-W1 is designed to combine two capabilities that are usually separated:
- Native multi-shot audiovisual co-generation with multimodal references.
- Native support for diverse real-time interaction signals through a streaming generation architecture.
This enables a more unified architecture than conventional cascaded systems that treat video generation, audio generation, and interaction as separate modules. It learns how user interaction signals, scene context, motion dynamics, audio events, multi-shot narrative generation, and visual references should jointly shape the next state of the generated world.
Capability 1: Native Audio-Visual Joint Generation
Vivix-W1 introduces native audio-visual generation. It can generate diverse voices and spatial sound effects grounded in visual events. The audio is tightly coordinated with the visuals, enabling smoother and more immersive storytelling. The model also supports a wide range of languages, while maintaining strong lip-sync, rhythm synchronization, and motion alignment.
Capabilities demonstrated by Vivix-W1 with approximately 30B active parameters, trained on a medium-scale dataset
Capability 2: Multimodal Reference Support
Vivix-W1 supports multimodal references, including images, videos, audio, text, and interaction context. This allows Vivix-W1 to generate worlds that are both controllable and dynamic.




Capability 3: Interactive Signal Understanding with Real-Time Streaming Generation
Vivix-W1 uses streaming generation. Instead of generating the entire video as a single closed clip, it generates the world continuously through fine-grained audiovisual segments. This fine-grained streaming approach allows Vivix-W1 to integrate diverse interaction signals directly into the generation process, so each audiovisual segment can be driven and adjusted in real time by user interaction.
Touchscreen Interaction
Motion- and Camera-Based Interaction
Capability 4: Native Multi-Shot Narrative Generation
Vivix-W1 introduces native multi-shot narrative generation. Unlike single-shot video generation systems, Vivix-W1 is designed to understand and generate coherent multi-shot sequences natively. It can coordinate scene cuts, camera motion, dialogue rhythm, character reactions, and multi-character interactions, allowing the generated content to feel like a coherently directed sequence rather than a collection of isolated clips.
Capability 5: Long-Horizon Streaming Generation for Continuous Narratives
To enable stable long-horizon real-time generation, we introduce Vivix-Turbo, a dedicated acceleration and consistency optimization technology for Vivix-W1. Vivix-Turbo is designed to push streaming world generation toward highly efficient real-time inference while maintaining long-horizon stability, multimodal consistency, and narrative continuity.
Vivix-Turbo optimizes the system across four key dimensions. First, it applies ultra-low-step distillation to reduce inference cost and make real-time streaming practical. Second, it jointly optimizes shot planning and temporal continuity over long sequences, allowing the model to generate extended multi-shot sequences without breaking narrative flow. Third, it reduces error accumulation over long streaming horizons, improving identity stability, scene consistency, camera coherence, and motion continuity. Finally, it strengthens multimodal reference consistency, ensuring that characters, voices, styles, environments, and behavioral patterns remain stable as the generated world continues to evolve.

Together, these optimizations allow Vivix-W1 to generate continuous, interaction-ready audio-visual streams with greater stability over extended durations. Instead of producing isolated clips, Vivix-W1 can maintain a continuously evolving generated world that remains coherent, responsive, and visually expressive over time.
Based on this capability, we can efficiently construct the following long-sequence audio and video streams within a time budget of less than one minute.
Limitations and Future Work
Validated at a medium training scale with approximately 30B active parameters, Vivix-W1 demonstrates a broad set of capabilities including multimodal understanding, real-time interaction, native audio-visual generation, multi-shot storytelling, and continuous streaming.
That said, Vivix-W1 is only the beginning. In certain complex scenarios, especially those requiring professional-grade visual fidelity, complex motion, or dense scene dynamics, there remains a gap compared with state-of-the-art clip-level video generation models trained on datasets containing billions of examples. These models benefit from significantly larger professional datasets and longer optimization cycles, especially for offline high-quality video generation.
We are now scaling Vivix-W1’s data engine toward tens of millions of hours of training data, expanding the coverage of cinematic scenes, multi-shot narratives, audio-visual events, character actions, spatial dynamics, and interaction patterns. As the data scale continues to grow, we expect Vivix-W1 to deliver significantly stronger stability, visual quality, complex motion, narrative coherence, and real-time interaction performance.
With full-scale training already underway for future versions, these results offer a compelling preview of where omni-modal interactive video is headed.