Artificial Analysis moves its video benchmark to 1080p and breaks models down by use case

Artificial Analysis has rebuilt its text-to-video benchmark with AA-Video-T2V v2.0 and a separate silent version. Models are evaluated at 1080p through human preference votes, then compared across ten use cases, ten capabilities, and eight visual styles.

One overall ranking is no longer enough

A model that performs well on live-action film is not necessarily the one that handles animated interfaces, on-screen text, or anime most effectively.

The new Artificial Analysis methodology is designed to expose those differences. AA-Video-T2V v2.0 evaluates text-to-video models that generate synchronized audio, while AA-Video-T2V-Silent v2.0 separately evaluates video without audio.

Each benchmark has its own Elo pool. Outputs with audio are therefore never directly compared with silent generations.

Videos are evaluated through blind pairwise comparisons. Reviewers watch two outputs generated from the same prompt without knowing which model produced either one, then select the result they prefer. For v2.0, Artificial Analysis reset the text-to-video ratings rather than carrying over votes collected under its previous methodology. 1,000 prompts across ten use cases and ten capabilities

AA-Video-T2V v2.0 uses 1,000 prompts, while the Silent benchmark uses 500. The prompt set is regularly refreshed as requests stop separating frontier models effectively.

Every prompt is tagged across two primary dimensions. Ten use cases cover areas including live-action film, marketing and advertising, retail, animation and gaming, architecture, social content, and UI/UX motion design.

Ten capabilities examine areas such as physics, human anatomy, camera control, text rendering, spatio-temporal consistency, audio synchronization, and dialogue with lip sync. Eight visual styles add another layer, including photorealism, 3D rendering, cartoon and anime, illustration, and stop-motion.

The overall benchmark samples evenly across use-case and capability combinations, making versatility a major component of the final score. Every video goes through the same evaluation framework

The target generation settings are 1080p, 24 frames per second, ten seconds, and a 16:9 aspect ratio. When a model cannot produce the exact setting, Artificial Analysis uses the closest supported option.

Videos are not upscaled for evaluation. Models unable to generate 1080p are shown at their highest native supported resolution, while outputs above 1080p are downscaled. Audio-enabled videos are leveled toward -18 LUFS without compression or remastering.

Artificial Analysis launched AA-Video-T2V v2.0 with more than 68,000 human preference votes across its 1,000 prompts. The Silent version started with more than 47,000 votes across 500 prompts. Wan 3.0 leads, but specialization varies considerably

The current leaderboard places Wan 3.0 first at 1157 Elo, followed by Utopai X, based on MiniMax H3, at 1150. Dreamina Seedance 2.5 follows at 1144, MiniMax H3 at 1139, and MiniMax H3 Max at 1134.

Confidence intervals overlap across several of those models, making a strict interpretation of adjacent positions less useful than the raw ordering might suggest.

Artificial Analysis' launch analysis also showed strong specialization. Wan 3.0 led ten of the twenty use-case and capability categories. Seedance 2.5 stood out in human anatomy and dialogue with lip sync, while Gemini Omni Flash 1.1 led audio synchronization and UI/UX & Motion Design.

Text rendering produced one of the widest performance spreads across the field, while camera control remained much more tightly grouped. Wan 3.0 was the exception in the initial analysis, leading the latter capability by 45 Elo. From $4.80 to $34.12 per minute near the top

Price varies substantially even among models with relatively close human preference scores. MiniMax H3, currently the highest-ranked open weights model, scores 1139 Elo at $4.80 per minute of 1080p video.

Dreamina Seedance 2.5 reaches 1144 Elo at $34.12 per minute.

Artificial Analysis calculates pricing from each provider's published cost to generate one minute of video under the benchmark's default generation settings.

Wan 3.0 currently costs $12 per minute while holding the top position. At the other end of the top ten, Agnes-Video-2.5 costs $2.40 per minute. The new benchmark therefore separates human preference, task specialization, and generation cost rather than reducing video model comparison to a single Elo score.