Phoenix-4.5 animates the upper body without slowing down the conversation.

Tavus launches Phoenix-4.5, a conversational avatar model that generates the face and upper body in real time, with an announced audio-to-video latency of 134 ms.

A face can follow a voice accurately while still seeming strangely motionless. The mouth says the right words, but the shoulders do not move, the posture remains frozen, and the character appears to be waiting for its next signal. Phoenix-4.5 attempts to correct that disconnect by extending animation to the upper body.

Introduced on September 10, 2026, Tavus’s new model jointly generates the face, head, neck, shoulders, torso, clothing, and visible surroundings. Movement no longer stops around the mouth or at the edge of the face: it is designed to follow the rhythm of speech and continue while the character is listening.

Phoenix-4.5 is not a complete conversational agent on its own. It represents the visual component of Tavus’s Conversational Video Interface. This pipeline also includes a perception system, a turn-taking module, speech recognition, a language model, and text-to-speech generation.

Tavus now uses “Face” to describe a character’s visual appearance. The term “PAL,” short for Personified Application Layer, refers to the configurable system that combines this face with a voice, knowledge, behavior, objectives, tools, and rules. A Conversation then connects the PAL to a participant through a WebRTC stream.

Phoenix operates at the end of this pipeline. It receives the audio produced by text-to-speech and translates that information into movement. The intelligence behind the response can come from another model, either provided by Tavus or selected by the developer. Phoenix-4.5 does not decide what the character should say. It primarily determines how the character should move while speaking or listening.

The previous generation used a square region centered on the face and head. That area was then composited onto prerecorded footage showing the rest of the body. This approach already made real-time facial animation possible, but it kept the torso in a largely fixed position.

Larger movements could also expose the boundary between the regenerated region and the original footage. A tilted head, a strand of hair, or an accessory could cross that boundary and create a visible break.

Phoenix-4.5 replaces this approach with a generative renderer rebuilt to process the entire frame in a single pass. The face is no longer added as a separate layer over a prerecorded body. Movement across the neck, shoulders, and torso is generated with it, without a mask or seam around the head.

Expression can therefore extend beyond facial features. The head follows the cadence of a sentence, the shoulders accompany a change in intonation, and posture shifts with the energy of the delivery. Tavus is not generating an entire body, however: the architecture described remains focused on a half-body representation.

The change also affects moments of silence. A person does not become perfectly still while listening. They adjust their gaze, subtly shift their posture, and react before speaking again. Phoenix-4 had already introduced active listening and continuous facial motion. The new version extends that behavior to the head and upper body.

The model analyzes both sides of the conversation. It receives the PAL’s audio as well as the other participant’s, allowing it to maintain a visible reaction while the user is speaking. The goal is to avoid an overly mechanical alternation between “speaking” and “idle” states.

Two models work together to produce the result. The first serves as the animator. A streaming diffusion Transformer converts both audio streams into compact motion representations while accounting for the intended identity, style, and emotion.

The system produces movement in successive blocks instead of waiting until the end of a sentence. Tavus says it distilled the model to reduce generation to a few sampling steps. This reduction makes it suitable for use during a live conversation.

The second model handles rendering. It uses a reference image and the motion received from the animator to produce the face and upper body. The architecture is based on a generative adversarial network and interprets movement relative to the reference instead of treating it as a set of absolute positions that must be reproduced identically.

This distinction is intended to make the same motion instruction compatible with different faces, hairstyles, and body types. Signals from the facial region also help coordinate the head, neck, and shoulders so that each area does not appear to move independently.

Tavus reports a latency of 134 ms between receiving audio and producing video, which it describes as a 25% advantage over any other available model. This figure does not represent the latency of an entire conversation.

Before Phoenix can begin its work, the system must detect the end or interruption of a speaking turn, transcribe the audio, prepare a response with a language model, and convert that response into speech. The 134 ms figure covers only the final stage connecting already-generated audio to its visual representation. It does not measure the time between the user’s question and the beginning of the response.

The publication also uses two slightly different figures. Its main announcement states 134 ms, while the technical section mentions latency below 130 ms and an end-to-end rate of 35 to 40 frames per second on machines described as slower. No public details specify the hardware, resolution used for each test, or how the measurements were aggregated.

The market comparison is equally difficult to reproduce. Tavus does not provide a list of the competing models tested, their settings, or a shared evaluation protocol. The claim that Phoenix-4.5 is 25% faster than any other system should therefore be treated as an internal measurement from the provider.

The other results primarily compare Phoenix-4 and Phoenix-4.5 on the same paired internal test set. The AVSR score used to evaluate audiovisual synchronization rises from 0.261 to 0.304. This improvement of approximately 16% corresponds to the figure Tavus highlights for better