Vivix-A1
Introducing Vivix-A1: The first real-time, human-centric interactive model for continuous perception, conversation, and action in an open world.
A truly human-centric model must go beyond talking-head and upper-body avatars. Digital interaction requires more than synchronizing lip movements with audio; it requires spatial awareness, physical interaction, and lifelike behavior in open-world environments.
Vivix-A1 supports spatial awareness, physical interaction, and lifelike behavior in an open world.
Vivix-A1 operates in an open world, where it can perform autonomous or instruction-guided actions while engaging in unstructured, real-time, full-duplex conversations.

Existing language-centric interaction pipelines typically rely on intermediate textual representations for reasoning and response generation, requiring user audio to be converted into language representations before reasoning can occur. As a result, they struggle to immediately respond to semantic information embedded in the audio stream itself, such as environmental sounds or acoustic events.
We introduce Vivix-A1, a full-duplex, human-centric model that replaces the conventional LLM–TTS pipeline with a native multimodal generation loop. As illustrated in Figure 1, Vivix-A1 can listen, reason, speak, and act continuously. Rather than routing every interaction through an intermediate language representation, the model grounds its outputs directly in raw multimodal signals. A Director Agent provides higher-level reasoning, response planning, and behavioral control, enabling more coherent, context-aware interactions over longer time horizons. This design supports more autonomous behavior in open-world environments while preserving low latency and full-duplex communication.
The table below compares Vivix-A1 with representative generative AI systems across key dimensions of real-time interaction.
| Seedance 2.0 | Kling 3.0 | TML-Interaction | LemonSlice | Tavus | Runway Avatar | Vivix-A1 | |
|---|---|---|---|---|---|---|---|
| Human-centric Interaction | |||||||
| Real-time Interaction | |||||||
| Environmental Audio Perception | |||||||
| Full-duplex Conversation | |||||||
| Open-world Autonomous Behavior | |||||||
| Multimodal Reference Support |
Based on product documentation and demos publicly available as of June 30, 2026.
Here is a deep dive into the architecture and capabilities that make Vivix-A1 possible.
A Native Multimodal-to-Video Architecture
Most voice-interactive systems and AI avatar platforms rely on turn-based paradigms and single-modality control. Vivix-A1 introduces a native multimodal, reference-conditioned architecture.
The model learns two complementary capabilities. First, it learns how to use multimodal information to drive visual character generation. To do this, we introduce an in-context learning mechanism that incorporates multimodal information directly into the streaming generation loop, allowing visual references, audio cues, and interaction history to jointly guide generation. Second, Vivix-A1 learns to natively interpret user interaction signals (e.g., environmental audio) and directly drive the real-time generation of streaming multimodal responses.

Expanding Streaming Generation for Ultra-Low Latency
Traditional visual generation models operate on large temporal blocks, typically producing 5-second or 10-second clips at a time. This clip-level generation paradigm is fundamentally incompatible with real-time interactive systems. Even if clip-level generation is implemented with real-time inference, it still introduces significant response latency.
To solve this, we redesigned the model architecture for streaming generation. Vivix-A1 produces output incrementally, one small multimodal segment at a time. Clip-level generation benefits from an overall temporal window: the model can jointly reason over several seconds of information to produce coherent visual and acoustic signals. In a streaming setting, preserving multimodal continuity, consistency, and stability throughout a sequence of incremental generations introduces unique challenges.
Sink tokens and reference anchors are commonly used to improve long-horizon stability. These mechanisms help the model maintain continuity over extended sequences. However, they also introduce a subtle trade-off: the preset signals that stabilize generation can constrain the model’s ability to produce diverse character appearances, visual dynamics, and open-ended behaviors.
To address these issues more effectively, we decouple the continuity signals from the diversity signals in multimodal generation. More specifically, we introduce local continuity priors between short multimodal segments, while using global sequence-level priors for long-horizon streaming generation. This design allows the model to maintain stable temporal coherence without over-constraining its ability to generate diverse visual behaviors, character dynamics, and open-ended actions.

Multimodal References as Behavioral Grounding
Vivix-A1 supports conditioning on multimodal references. These references can define a character’s identity, appearance, voice, environment, and interactive objects. Unlike conventional reference-based generation, where references mainly serve as initial anchors for identity, appearance, or style, Vivix-A1 treats them as persistent behavioral grounding signals. Throughout continuous streaming generation, these multimodal references can dynamically modulate the generated multimodal signals over time—across voice, motion, and interactive behavior. This design enables dynamic control over the generated multimodal stream, allowing users to modify the generated world during an ongoing interaction. For example, users can provide a product reference and ask the character to pick it up and present it.
Vivix-A1’s subsequent behavior can be dynamically shaped by objects introduced into the scene.
Beyond Avatars: Listen, Speak, and Act
Many current avatar systems excel at matching lip movements to audio, but typically keep the character in a fixed pose or a tightly constrained action space.
Vivix-A1 extends beyond this paradigm. Rather than remaining fixed in place, it can respond directly to instructions and perform a wider range of actions in open-world environments.
- Spatial Exploration: Vivix-A1 can walk, turn, and navigate within an open-world scene while conversing.
- Human-Object Interaction: It can interact with objects in the generated world, including existing objects in the current scene as well as new entities introduced dynamically during streaming generation.
- Unconstrained Poses: It moves beyond the fixed, desk-bound "news anchor" setup.
Its native architecture enables Vivix-A1 to generate facial expressions and full-body motion while interacting with its environment in a context-aware manner.
Vivix-A1: Listening and Environmental Awareness
Vivix-A1 Reacts to Thunder
Vivix-A1 Reacts to Flirtatious Conversation
Vivix-A1 Hears "Hi"
Vivix-A1 Reacts to Strong Wind
Vivix-A1: Expressive Speech
Vivix-A1 Expresses Sadness
Vivix-A1 Expresses Anger
Vivix-A1 Expresses Joy
Vivix-A1 Expresses Surprise
Vivix-A1: Open-World Free Behavior
Multidimensional Joint Distillation for Ultra-Efficient Inference
Distillation is a common optimization stage in today's state-of-the-art real-time avatar systems. Unlike strategies designed for fixed visual composition and constrained actions, Vivix-A1 operates in a much more open generation space, which makes stability and consistency significantly more challenging. We therefore introduce a Multidimensional Joint Distillation strategy that progressively optimizes the model across three complementary dimensions:

- First, we optimize short-horizon generation for few-step inference. While high-quality video diffusion systems often rely on 4 to 8 inference steps to preserve generation quality, our distillation approach reduces Vivix-A1’s inference process to only 2 steps while maintaining reliable visual expressiveness and motion quality.
- Second, we continue to optimize long-horizon multimodal stability. A major challenge in open-world generation is the accumulation of errors over time, including background warping and visual drift. Naively applying existing distillation methods may preserve consistency, but often at the cost of motion quality and visual diversity. To solve this, we introduce End-to-End Self-Distribution Consistency Optimization. In practice, we leverage the teacher model’s diverse motion behaviors to guide the student model in preserving motion diversity while preventing long-term multimodal drift.
- Third, we optimize multimodal performance across multiple representation spaces. In our early experiments, optimization in the VAE-compressed latent space alone was insufficient: it failed to capture fine-grained local dynamics and visual details and was less sensitive to accumulated long-horizon errors. Based on this exploration, we constructed a joint distribution optimization strategy across both latent and raw data spaces. This dual-space distillation improves local dynamics while reducing long-horizon drift.
Model Evaluation
We built Vivix-Bench, a benchmark containing 1,000 carefully curated samples, to compare Vivix-A1 with representative streaming avatar systems. To ensure a fair and rigorous evaluation, we selected publicly accessible models for comparison and standardized the input conditions across all systems, enabling a more controlled comparison of model capabilities. The benchmark covers a wide range of image styles, including photorealistic, artistic, anime, and stylized content, and features both human and non-human subjects. It further evaluates models across diverse motion instructions, open-ended environment interactions, expressive emotional states, and complex subject-object dynamics. We use pairwise human evaluation based on a Good/Same/Bad (GSB) protocol.



Experience Vivix-A1 Today
We believe that the future of interactive AI does not lie in forcing all interactions through disconnected modality-specific pipelines. Fluid interaction requires a native multimodal foundation. With Vivix-A1, we have taken an initial step toward full-duplex multimodal understanding, multimodal streaming generation, open-world action, and computational efficiency. We will continue scaling this system to explore deeper capabilities and architectural breakthroughs. We believe these methods can ultimately provide the foundation for training more capable interactive agents in open-ended simulated worlds.
Limitations
Open-world full-duplex character generation remains an extremely challenging problem. Unlike fixed talking-head avatars, it requires the model to jointly generate speech, body motion, scene dynamics, and object-level interactions in a continuous streaming loop, where small temporal errors can accumulate over long-horizon interactions.
With Vivix-A1, we have taken a first step toward addressing this challenge through a native multimodal architecture and dedicated optimizations for streaming generation. However, in complex scenarios involving rapid motion, spatial navigation, or fine-grained human-object interaction, the current model may still produce motion inconsistencies, interaction drift, or imperfect physical grounding.
In future versions, we expect to mitigate these issues by training on larger datasets of long-horizon interactions, broadening coverage of human-object interactions, and strengthening multidimensional consistency optimization across motion, visual, and audiovisual representations.