A selfie and a voice clip, and an avatar speaks your texts in Google Vids

Google Vids combines Gemini Omni and digital voice doubles to automate editing. How text-based editing changes your projects.

Gemini Omni and personal avatars are joining Google Vids, the video creation tool integrated into Workspace. The former generates and edits clips from a description in natural language. You start with a text prompt, add visual references, photos, or rough sketches as needed, and Omni blends these inputs to produce a video. Google had already opened Veo 3.1 to all Vids users in February.

Editing is also done through conversation. Whether on a shot generated by Omni or on a sequence filmed with a phone, you describe the desired change to replace a background, correct the lighting, or add an effect. Edits are made step-by-step, without starting from scratch.

Personal avatars open up another path. From a selfie and a short voice recording, the tool creates a digital double that replicates the user's face and voice. You then type the text to be spoken, and the avatar delivers the message without you having to go in front of the camera.

Both features remain reserved for Google AI Pro and Ultra subscribers and Workspace business customers. Access to avatars is currently limited to certain regions and to adults; each avatar remains linked to the Google account and solely to the appearance of its owner. Each generated clip embeds an invisible SynthID watermark, meant to certify that a video was produced by AI.