Ten seconds of rushes and the double is ready
Captions' Avatar X model generates a complete digital double from a short recording, promising speed for video production.
Mirage, Captions’ model lab, is pushing Avatar X, its next generation of video avatars. The claimed shift lies in the production method. While the common practice consists of separately producing the image, voice, and movement, and then synchronizing them, the model generates the entire performance in one go. The company expects this to yield superior consistency in the details that usually betray assembly, particularly lip alignment and gaze direction.
Two known flaws of the genre are being targeted. First, identity preservation—that likeness that degrades as a video gets longer, and which Mirage claims to maintain over continuous generations. Second, micro-expressions—those brief movements that make a face look alive rather than rehearsed, such as laughs, yawns, or grimaces. The generated expressions are not confined to the lower face and extend to the body, all driven by the audio. Both vertical and horizontal formats are supported.
The most concrete point concerns the input. According to the publisher, ten seconds of video are enough to capture face, voice, and identity to create a digital double. Avatar X now powers all of Captions’ avatars and doubles, where users choose a figure, write a script, and start the generation, with the existing library remaining accessible and custom avatars able to be produced via prompt.
Two caveats accompany this reading. The comparison with competing models, which is heavily emphasized, comes from the company itself and is not accompanied by any documented protocol or third-party evaluation. And nothing in what has been published addresses the consent of the cloned person, the watermarking of the produced videos, or the safeguards surrounding a double made from a ten-second recording.