Synthesia turns video production into a conversation

Synthesia Assistant creates and edits videos from a prompt, document, or URL, featuring Express-3 avatars, brand identity, and access to the full editor.

A prompt describes a welcome video for a new hire. Synthesia Assistant uses it to devise a structure, write the scenes, select an avatar, and prepare the visuals. Further changes are handled in the same conversation: shorten the script, adjust the tone, replace a sequence, or reorganize the presentation.

With this new Assistant, Synthesia aims to bring the different stages of its platform into a conversational interface. The service does more than generate a video from a few lines of text. It organizes the project before production, builds a first cut, and then applies requests made after the video has been created.

Users can start with a basic description, an existing script, a webpage, or a file. The Synthesia Assistant product page lists support for PPTX, PDF, DOCX, and TXT formats. A training document, internal procedure, or existing presentation can therefore serve as source material.

Assistant first proposes a plan outlining the structure, number of scenes, and overall tone. This plan can be edited before the video is produced. Users can rewrite sections, move scenes, select a visual template, or choose the avatar that will present the content.

This approach addresses a common limitation of automated generators. A vague prompt may lead to a polished result that is poorly structured, too long, or disconnected from the original goal. By presenting its plan first, Synthesia gives users a chance to correct the direction before the scenes are assembled.

The system then prepares the script, sequences, visual treatment, B-roll, and motion graphics. It can also apply the company’s Brand Kit and custom templates. Synthesia says these guidelines cover avatars, illustrations, animations, and additional footage, rather than just the graphic frame surrounding each slide.

Assistant is not a new standalone video model comparable to services specializing in cinematic clips. It primarily acts as an orchestration layer on top of Synthesia’s existing features. Its role is to coordinate writing, staging, media selection, avatars, and editing.

The examples on the public page nevertheless extend beyond traditional corporate presentations. They include a dog running through a park, explorers inside an ice cave, and a neon-lit alley in heavy rain. These demonstrations show the broader range of available visuals, but Synthesia does not promise that a long, continuous scene will be generated in a single pass. The result is still organized as a sequence of scenes inside its editor.

A company can import an approved script without requesting a rewrite. Synthesia says the wording will then be preserved exactly. This option is relevant to fields where every sentence must be approved before publication, including compliance, regulatory training, and legal communications.

Preserving the text does not remove the need to review the final result. An unchanged sentence could still be paired with an unsuitable visual, delivered with an unnatural intonation, or placed in a scene that alters how it is understood. Approval therefore needs to cover the complete video, not just the script.

Once the first cut has been produced, further requests are handled through the chat built into the editor. Users can shorten a passage, change its register, rewrite a scene, or ask for new visuals. Each change is applied to the existing project without requiring the video to be rebuilt from scratch.

This continuity is the product’s most significant change. Earlier generation assistants focused mainly on producing an initial draft. Corrections then had to be made with conventional tools, sometimes one scene at a time. Synthesia now wants the conversation to continue throughout the revision process.

The full editor remains available at any stage. A team can begin with Assistant, then manually adjust the timeline, scripts, media, or scene layouts. The chat does not entirely replace the existing interface or lock users into its initial choices.

This combination is aimed primarily at onboarding videos, internal training, sales presentations, and employee communications. In these settings, production speed often matters more than an original visual treatment. Approved templates and brand guidelines can also reduce inconsistencies between videos created by different teams.

The launch also introduces Express-3, Synthesia’s new generation of avatars. The company claims it offers better lip-syncing, gestures more closely connected to the spoken words, and body movements that adapt to the script. Facial expressions are also expected to change according to the tone of the text.

Synthesia describes Express-3 as its first avatar generation to outperform Express-2 in direct comparisons involving lip-syncing, gestures, and body movement. The launch page does not provide detailed scores, the number of evaluators, or a complete testing protocol. The claimed improvement therefore remains difficult to quantify from the available information.

The quality of an avatar is not determined by the accuracy of its mouth movements alone. Voice pacing, pauses, eye movement, posture, and the relationship between gestures and speech also affect credibility. An improvement visible in a short demonstration may be less consistent during a longer presentation or in a script containing unusual terminology.

Express-3 is integrated into Assistant. Once a presenter has been selected, the service manages their placement and performance across the different scenes. Users can still change the avatar or revise the layout in the editor.

Synthesia is also launching Avatar Builder, a tool for creating a character from a text prompt, preset, or reference image. The available styles range from cartoons and anime to photorealistic renderings. The character can then be given a name and voice before being used in a video.

These generated characters should be distinguished from the Personal Avatars and Studio Avatars already available on the platform. Those products seek to reproduce a real person using a