On one minute of generated film, a benchmark puts Runway in the lead
The Physion-Arc 1.0 benchmark evaluates six video agents over a one-minute duration. Runway Agent 2.0 leads the ranking but still struggles with chapter transitions.
A new benchmark from Physion Labs, Physion-Arc 1.0, tackles one-minute videos produced by video agents, whereas its predecessor Atlas 1.0 was limited to clips of just a few seconds. The stated challenge goes beyond technical correction: a video can follow its script while still being poorly directed, flat, or visually inconsistent.
The protocol is organized around three dimensions, broken down into sixteen metrics: narrative consistency, cinematic language, and production quality. To pinpoint flaws at the right level, the video is divided into chapters and the script into beats—the finest unit marking a change in action or intention. Human evaluators score objective criteria on a grid, while subjective qualities—pacing, emotional impact, cinematic taste—are assessed through blind A/B comparisons aggregated in an arena-style format.
One hundred scripts from T2F-Bench were submitted to six commercial agents—Runway Agent 2.0, Utopai PAI 2.0, MiniMax Hub, Luma Creative Agents, TapNow Agent, and Kling Canvas Agent—resulting in six hundred one-minute videos. According to Physion Labs, Runway Agent 2.0 ranks first overall with 87 out of 100 points and dominates all three dimensions, with a clear lead on subjective metrics. One key takeaway emerges: consistency holds up well within a chapter but degrades from one chapter to the next, with character identity being the first to slip.
The team also notes that none of the six agents operated without intervention, due to stalled generations and manual reassembly. They link the project to a broader effort: world critics capable of judging whether a video maintains its consistency, pacing, and meaning over time, extending their work on the evaluation of World Models.