OpenArt Arena ranks models based on creative briefs
OpenArt Arena ranks image and video models across twelve creative use cases. Its blind evaluations involve 1,031 professionals, but remain limited to the platform's catalog.
Seedream 5.0 Pro leads the overall image ranking, but GPT Image 2 moves ahead in graphic design and image editing. In video, Seedance 2.5 holds the top overall position, while Wan 3.0 takes the lead in video editing. These differences capture the idea behind OpenArt Arena: the best model depends on the work being requested.
The ranking published by OpenArt divides image and video into 12 leaderboards. Instead of assigning a single score meant to cover every use case, it separates advertising, filmmaking, animation, motion graphics, graphic design, e-commerce, image editing, video editing, and lip-syncing.
This structure is meant to bring testing closer to the decisions made during production. An agency concerned with preserving a logo will not necessarily select the same model as a filmmaker seeking continuity across several shots. A localized edit also requires different strengths from an original creation generated from text.
The first version of the index, dated September 15, 2026, includes seven video leaderboards and five for image models. The overall tables measure capabilities shared across multiple use cases. The remaining rankings apply criteria specific to an industry or operation.
Five dimensions make up the overall video score: prompt adherence, subjective aesthetics, convincing physics and motion, visual consistency, and audio quality. Each accounts for 20%.
For images, four criteria receive equal weighting: prompt adherence, subjective aesthetics, reference adherence, and the ability to offer a creative interpretation. OpenArt explicitly treats “aesthetics” as a subjective judgment rather than presenting it as an objectively measurable property.
The specialized leaderboards then adjust this framework. For video advertising, evaluators examine text and logos, brand consistency, product realism, and whether the product makes sense within the scene. The film ranking adds camera and lighting control, acting performance, continuity between shots, and understanding of the requested style.
Graphic design gives separate weight to typography, layout, and adherence to the visual direction. Image editing prioritizes the precision of a requested change and the preservation of elements that were not supposed to be altered. Lip-sync tests distinguish between speech, singing, multiple-character scenes, and different languages.
Each comparison begins with the same brief being sent to the selected models. OpenArt generates one output per model according to a fixed rule and says no result is manually selected from several attempts. Two anonymized creations are then displayed side by side.
Evaluators do not answer a broad question such as, “Which one do you prefer?” They vote on a named criterion. A comparison might ask which video follows the instruction more closely, which one preserves a face more faithfully, or which one produces the most convincing movement. Model identities remain hidden, and their placement on the left or right is randomized.
Separating the criteria addresses a common problem with preference voting. A spectacular result may make a strong first impression while ignoring part of the brief. By isolating fidelity, continuity, or text rendering, the Arena attempts to measure what makes an output usable, not simply which one has the strongest immediate appeal.
The published scores come from 1,031 participants. The panel presented by OpenArt includes 28 members of a Creative Expert Council and 1,003 “tastemakers.” The first group includes filmmakers, producers, creative directors, educators, and specialists in AI-assisted creation. The second includes agency professionals, studios, students, creators, and members of the OpenArt community.
Their votes do not carry equal weight. A council member’s vote counts three times as much as a tastemaker’s. This gives the group selected by OpenArt greater influence, even though most of the voting volume comes from the broader panel.
Several controls are designed to identify inconsistent voting. Between 2% and 3% of the comparisons assigned to certain judges repeat a pair they have already reviewed, with the sides reversed. These responses measure participant consistency and are not included in the final calculation. No judge may contribute an excessive share of the votes within their tier.
OpenArt also checks the rankings by temporarily changing the weight assigned to experts, removing the judge who completed the highest number of comparisons, and altering the statistical sampling used to calculate uncertainty. If the order changes during these tests, publication of the leaderboard is paused for review.
The methodology uses a Bradley-Terry statistical model. It estimates the relative strength of each competitor from pairwise matchups. Consistently beating a highly ranked model has a greater effect on the result than winning against one near the bottom of the table.
A reference model anchors each leaderboard at 1,000 points. The displayed figures therefore resemble Elo scores, but they are neither scores out of 1,000 nor success rates. A result of 1,125 indicates a model’s relative position within the relevant leaderboard.
Scores from different categories should not be compared directly. A result of 1,050 in advertising and the same figure in film come from different criteria, weightings, and matchups. Every score also carries a 95% confidence interval. When the intervals of two models overlap, OpenArt considers their performance statistically indistinguishable despite the order shown in the ranking.
The first overall video leaderboard places Seedance 2.5 at the top with 1,125 points, ahead of Wan 3.0 at 1,047 and Seedance 2.0 at 1,044. Around 3,800 votes were collected for each model in this table. Seedance 2.5 also receives the highest individual score across all five general criteria.
The hierarchy becomes much tighter when the task changes. For video editing, Wan 3.0 reaches 1,080 points, compared with 1,079 for Seedance 2.5 and 1,068 for Google Omni