Vivix-W1 lets interactions rewrite the world in real time

Vivix-W1 generates a continuous audiovisual stream that text, voice, images, gestures, and camera movements can alter while it is playing.

A scene is already unfolding when a new image adds an object, a click redirects a character’s attention, or a voice instruction changes the action. Vivix-W1 does not require the entire generation to restart: each interaction affects the part of the world that does not yet exist.

Vivix describes W1 as a streaming-native multimodal model. Here, the term refers to a system designed to produce a continuous audiovisual stream rather than a closed video calculated in full from an initial prompt.

The difference lies less in the formats it accepts than in when those formats can be introduced. An image can establish the setting at launch, while another can add an object during playback. Text can define the initial situation or redirect the story several moments later. Voice, video, and motion controls also remain available after generation has started.

Vivix summarizes this structure with a simple principle: a prompt defines the content, while an interaction changes its future. Commands do not erase what has already been shown. They guide the upcoming segments from the state the world has reached when the input is received.

W1 generates a succession of short audiovisual segments. Each completed section becomes part of the context used to prepare the next one. The system combines it with new instructions, multimodal references, and the interaction history to continue the stream in an updated direction.

This approach makes W1 closer to an interactive session than a conventional video generator. The user no longer submits a request and simply waits for an output. They can intervene as it unfolds, much as they would in a game, narrative experience, or previsualization environment.

Controls include text, voice, clicks, touchscreen interactions, live image references, motion, and camera direction. The API documentation defines events that can contain text, references, images, or audio signals. These events are applied to the portion of the session that has not yet been generated.

Clicking an object can therefore become a narrative instruction. Movement can change the viewpoint. A photograph can introduce an appearance, location, or accessory. A spoken command can ask a character to stop, turn around, or react differently.

This flexibility does not mean that every interaction will produce an immediate or perfectly predictable result. The model must receive the signal, interpret it, add it to the context, and generate a continuation compatible with what came before. Vivix claims responsiveness at the scale of seconds but does not provide complete measurements showing the delay between an action and its visible consequence.

The presentation occasionally refers to the millisecond-level latency requirements of the underlying architecture. This describes the speed required for some internal processing, not a guarantee that an entire world will respond within a few milliseconds. The documentation does not yet include median latency, higher-percentile results, or variations under heavier demand.

W1 also generates visuals and sound as part of the same stream. It can produce voices, dialogue, ambience, and spatial sound effects tied to visible events. This joint generation is intended to reduce the discrepancies that appear when video, speech, and soundtracks are created separately and combined afterward.

The model supports sequences composed of multiple shots. It can choose a new viewpoint, move the camera, arrange a cut, and continue a conversation in the following shot. Vivix is therefore assigning the model some of the directing work that would normally be handled by a separate production pipeline.

This autonomy does not remove human direction. Users can still provide visual guidance, references, or new instructions. W1 instead attempts to handle the intermediate decisions required to keep a situation moving without waiting for a command at every step.

Characters, objects, environments, voices, and visual styles must remain stable enough across segments for the result to feel like one continuous world. That consistency is one of the system’s central challenges. A minor deviation repeated in every new section can eventually alter a face, move an object, or gradually reorganize a location.

Vivix pairs W1 with a technology called Vivix-Turbo. It uses ultra-low-step distillation to accelerate generation, while also addressing shot planning, temporal continuity, and the accumulation of errors throughout the stream.

The company says this approach improves the stability of identities, scenes, camera movements, and multimodal references across longer sequences. Its report explains these principles but does not provide detailed numerical results, a reproducible testing protocol, or a comprehensive comparison with competing systems.

The model presented by Vivix has approximately 30 billion active parameters. This figure does not necessarily reveal its total size, since an architecture may use only part of its parameters for each operation. Vivix does not disclose enough information to determine W1’s exact structure, hardware requirements, or the conditions under which it reaches real-time performance.

The current version was validated at what the company calls a medium training scale. Vivix says it has already assembled tens of millions of hours of data for its next development phase. That volume primarily concerns future versions and does not automatically describe the material used to train the model being presented today.

The origin, composition, and licensing conditions of this material are not detailed in the public report. A collection combining video, voices, music, movement, and narrative situations raises specific questions about rights, consent, and the representation of individuals.

Vivix also acknowledges several limitations. W1 remains behind leading offline clip-generation models in dense scenes, complex motion, and productions requiring professional visual fidelity. Those systems can spend more time on an individual