Vivix-W1: A Glimpse into a Streaming-Native Multimodal Model for Continuous Generation Shaped by Interaction
Most video models treat text, images, audio, and other modalities as conditions that define the video upfront, largely fixing its course before generation begins.
Vivix-W1 explores a different paradigm: interaction itself becomes a live condition that continuously shapes what happens next.
A prompt defines the content. An interaction changes its future.
Vivix-W1 natively supports text, images, audio, video, and other modalities as references for generating rich audio-visual worlds. It also allows new interactions to enter throughout generation and continuously shape what happens next.
This interaction space extends far beyond the navigation and camera controls commonly associated with today’s world models. It can encompass real-time touch input, camera and character movement, voice, text, images, video, and other control signals. And this is only the beginning. We are actively expanding the ways people can interact with Vivix-W1, with more forms of input still to come.
Validated at a medium training scale with approximately 30B active parameters, Vivix-W1 delivers a broad set of production-ready capabilities, including multimodal understanding, real-time interaction, native audio-visual generation, multi-shot storytelling, and continuous streaming.
With tens of millions of hours of training data already assembled for the next stage of scaling, the upcoming version is expected to push these capabilities significantly further. These results offer a compelling preview of where omni-modal interactive video is headed.
See Vivix-W1 in Action
Beyond the Prompt. Into Open-Ended Interaction.
In Vivix-W1, the same modality can define the starting point or reshape the world later.
An image can establish the initial scene—or introduce a new object midway through it. Text can set the starting context—or redirect how the world evolves after generation has begun. Audio and video can likewise serve as initial references or become new inputs within an ongoing stream.
Interactive world models have already shown how navigation and camera controls can influence a generated environment in real time. Vivix-W1 extends this idea into a much broader interaction space—spanning text and voice instructions, images and video, touch and click interactions, camera and motion controls, object selection, and real-time changes to objects, environments, and events within the world.
What matters is not only which modalities a model accepts, but whether they can remain active throughout generation—continuously shaping an evolving audio-visual world.

Vivix-W1 combines multimodal references with live interaction signals, using each newly generated audio and video segment as context for what comes next.
More Ways to Interact. Every Input Shapes the World.
Vivix-W1 supports a broad and continuously expanding range of interactions. New inputs can enter while the audio-visual world is already unfolding, shaping what happens next without restarting the generation process.
This goes far beyond the navigation and camera controls commonly associated with interactive world models. Users can touch what they see, move through the environment, direct the camera or action, provide text, voice, images, or video, and introduce changes to the world in real time. Each interaction becomes part of the ongoing context, allowing the world to respond from its current state and continue along a new path.
Voice & Text Input
Mouse Interaction
Touchscreen Interaction
Live Image Reference
Motion & Camera Control
A Director Built into the Model
Vivix-W1 natively directs story, cinematography, and sound together. As events unfold, it generates and organizes multi-shot sequences, deciding when to cut, shift viewpoints, move the camera, and advance the story while preserving the rhythm of dialogue and action.
Voices, dialogue, environmental ambience, spatial sound effects, music, motion, and visual events are generated together in one coordinated stream. This native integration gives each shot the right perspective, timing, and sound, creating sequences that feel coherent, intentional, and cinematic.
Native multi-shot storytelling with dynamic scene cuts and synchronized dialogue, action, and sound.
Native Multimodal Understanding and Reference Support
Vivix-W1 natively understands text, images, audio, video, and combinations of modalities as references for generation. These inputs can define the appearance, environment, visual style, motion, sound, and broader context of an audio-visual world.
References are not limited to the beginning of generation. They can also enter while the world is already unfolding. An image can introduce a new object or environment, a video can provide motion or visual context, and audio can add a voice or acoustic reference. Vivix-W1 carries these signals into the continuing stream, allowing new information to become part of the existing world rather than beginning a separate clip.
Case 1



Case 2



Multiple references work together to define the people, objects, environments, motion, and visual style of the generated world.
Experience Vivix-W1
Vivix-W1 is our first step toward audio-visual worlds that do more than play back—worlds that remain open to interaction, evolve in real time, and respond as new inputs shape what happens next.
Explore how Vivix-W1 can enable a new generation of interactive storytelling, entertainment, simulation, and creative experiences.
