Griffin listens, watches, and speaks at the same time during a video call
Tavus unites audiovisual perception, continuous conversation management, and video generation in Griffin, its first Human Interaction Model. Griffin-Lite generates a complete 720p scene and operates in full-duplex, with conversational decisions made at sub-second intervals.
A Conversation That No Longer Works in Turns
A pause, a glance, or an interruption can modify the response even while it is already underway. With Griffin, Tavus aims to break away from the usual sequential operation of video agents, where speech recognition, reasoning, and synthesis occur one after the other.
Griffin-Lite is presented as the company's first Human Interaction Model, or HIM. The system continuously receives audio and video from its interlocutor while producing its own voice, expressions, and movements.
It can thus nod while a person is speaking, interrupt or accept being interrupted, wait during a silence, or react to an item shown in front of the camera. In demonstrations published by Tavus, Griffin, for example, tracks the manipulation of a Rubik's Cube, reacts to the appearance of an object during a story, or guides a person using a soldering iron. Continuous Decision-Making
The architecture separates two major functions. A Continuous Conversational Modeling engine analyzes audio and video at sub-second intervals, then decides what to say, when to intervene, and how to behave.
These decisions then feed into an Audio-Visual Generation engine responsible for producing streaming speech and video. The transmitted controls also cover emotional tone, gaze, facial expression, and gestures.
Perception remains active while Griffin is speaking. A change observed in the middle of a sentence can therefore modify the rest of the response without waiting for a new turn in the conversation. The Entire Scene Is Generated
Unlike an avatar animated around the face, Griffin generates the entire image from a single reference. The body, hands, nearby objects, shadows, and background can evolve during the exchange.
The video generator produces 720p in 320 ms segments. An 8× temporal compression reduces eight frames to one latent, while autoregressive generation uses three diffusion steps per latent. Tavus states it has combined Distribution Matching Distillation, teacher forcing, and then Self-Forcing to maintain stability during long conversations.
On the audio side, Griffin-Lite uses Tavec, a convolutional autoencoder working at 48 kHz. The representation retains 40 values per frame at a rate of 100 frames per second, with a causal decoder designed for streaming. NVIDIA Places Griffin-Lite at the Top of VideoFDB
Results from NVIDIA's VideoFDB benchmark provide an external evaluation to Tavus's own tests.
On the perception component, Griffin-Lite achieves an overall score of 3.73 out of 5, compared to 3.44 for MiniCPM-o 4.5 in its audio-only configuration and 3.17 for Gemini 2.5 Flash Native in audio-video. The human baseline reaches 4.20. NVIDIA also measures a conversational timing alignment of 73.8% for Griffin-Lite.
Tavus also reports a score of 3.83 on the generation component of VideoFDB, compared to 3.92 for the human baseline and 2.80 for Gemini 2.5 paired with Anam. 26 Out of 54 People Thought They Were Talking to a Human
The 48% figure put forward by Tavus comes from a separate experiment conducted by the company. 54 participants were recruited via an independent research platform and placed in a one-minute video call with Griffin-Lite without knowing that their interlocutor was AI-generated.
At the end of the experiment, 26 out of 54 participants, or 48%, stated they thought they had spoken to a real person. With Tavus's previous system, combining Phoenix-4.5, Sparrow-2, and Raven-1, the same protocol had obtained 1 positive response out of 41 participants, or 2.4%.
Tavus describes this result as the first passing of a "video Turing test." However, this is not a standardized protocol or an independent benchmark comparable to VideoFDB, but rather a study designed and published by the company itself. Still a Limited Release
Griffin's capabilities also directly raise the question of identifying one's interlocutor. Tavus acknowledges that a system realistic enough to be mistaken for a person can also be used to deceive the user.
For this reason, Griffin-Lite remains limited to a selection of testers as part of a research preview. The company is working on disclosure mechanisms and indicates it wants to address these safety issues before opening Griffin more widely.