HappyHorse 1.1: The New Version of Alibaba's Video Model

Alibaba releases HappyHorse 1.1, an upgraded video generation model by its ATH unit featuring fluid motion, multi-image referencing, and robust prompt following.

Alibaba is pushing a new version of HappyHorse, its video generation model developed by the ATH unit. Version 1.1 builds upon the foundations of 1.0 and tightens four aspects of video production.

The name already carries a reputation: first noticed in the spring after an anonymous appearance on a public ranking, the model had topped blind tests before Alibaba declared itself its author.

The first area of focus is movement, which Alibaba says is more fluid and expansive, thanks to reworked modeling and improved frame-to-frame continuity. Second, multi-image referencing, known as R2V, aims to preserve the identity of a single subject across multiple character sheets without confusing them, by separating character images from background images. Next comes instruction following, which is more robust for long prompts: a single instruction can describe a sequence of scenes, with the model itself distributing the duration and camera cuts. The last focuses on image quality, with less artificial textures, softened skin without a plastic effect, and refined close-ups for micro-dramas and advertising. Audio, meanwhile, continues to be generated simultaneously with the image, including timed dialogue and lip-sync, as in the previous version. The model is used via API.