ElevenCreative brings together voice and face in its new Avatars
ElevenLabs voices and lip-sync models unite in ElevenCreative Avatars, a new single-interface tool for generating talking videos from text or images.
ElevenCreative introduces Avatars, an entry point for producing talking videos by combining ElevenLabs voices with lip-sync models within a single interface. Where previously one had to navigate through multiple tabs and choose a lip-sync provider, now a single screen is sufficient: you choose an avatar, write a script, select a voice, and launch the generation. With Text-to-Speech integrated into the input area, the voice and synchronized video are generated together, without a separate audio export.
An avatar is a persistent visual identity, human or non-human, created from reference images or a text prompt, with a modifiable default voice. You can create variations, angles, outfits, or backgrounds, while maintaining the same identity, then reuse the avatar in any generation from its Assets. The source images serve as a character sheet to maintain consistency across longer content.
For automation, an Avatar node has been introduced in Flows: the same pipeline can generate a script, a voiceover, and a video, then run in batches for multiple products, languages, or taglines. A library of ready-to-use avatars completes the offering. The feature is available on all paid plans.