HeyGen launches Avatar V, a video avatar generation system with persistent identity consistency

HeyGen launches Avatar V, a video generator using a Diffusion Transformer to keep identity consistency across clips, beating Kling O3 Pro and Veo 3.1.

HeyGen unveils Avatar V, a new version of its video avatar generation engine. The system captures an individual's identity from a 15-second reference sequence and preserves it across all generated videos, regardless of their duration. Outfit, background, appearance: visual variables can be freely modified without altering the subject's unique characteristics, whether it's their gestures, lip-sync, or micro-expressions. Technically, the model relies on a Diffusion Transformer conditioned on the full token sequence of the reference video, without identity compression. A five-stage training pipeline integrates fine-tuning, distillation, and RLHF alignment. On published benchmarks, Avatar V outperforms Kling O3 Pro, Veo 3.1, OmniHuman 1.5, and Seedance 2.0 in lip-sync, facial similarity, and six dimensions of human evaluation. The full technical report is available on the HeyGen Research website.