P-Video-2-Pro accelerates MiniMax H3 by reducing its scope
Pruna launches P-Video-2-Pro, an optimized version of MiniMax H3 that produces 5-to-15-second videos with sound from text or an image.
A five-second video generated in less time than it takes to watch. That is the main promise of P-Video-2-Pro, Pruna AI’s new model based on MiniMax H3 and designed to rapidly generate video with sound.
P-Video-2-Pro turns a text prompt or starting image into video. It can also accept a second image defining the desired final frame. Clips range from 5 to 15 seconds, are generated at 24 frames per second, and come in either 480p or 768p.
Pruna reports a generation time of approximately 1.99 seconds for a five-second 480p video in Speed mode. The same duration at 768p reportedly takes about 4.3 seconds. In both cases, the generation could therefore finish before the resulting clip has finished playing.
That speed does not mean Pruna trained an entirely new video model. P-Video-2-Pro is based on MiniMax H3, released in late July 2026. The German company adapted that foundation to reduce execution time, cost, and interface complexity.
MiniMax H3 was originally designed as a much broader multimodal system. It can accept text, images, video, and audio, then generate clips of up to 15 seconds at 2K with stereo sound. It also supports instruction-guided editing, motion transfer, and the simultaneous use of multiple references.
P-Video-2-Pro does not expose all those capabilities. Its interface is limited to text, a starting image, and an optional ending image. It accepts no audio track, reference video, or footage to edit. Its maximum 768p resolution also falls below the 2K supported by the original model.
This narrower scope likely contributes to the speed gains. The service focuses on two tasks: generating a video from a description or animating an existing image. Pruna has not disclosed the exact process applied to H3 to create this version.
The company develops a range of optimization techniques that include lower-precision computation, distillation, pruning, compilation, and caching. No technical report specifies which of these are used here, in what proportions, or whether the MiniMax H3 weights underwent additional post-training.
The “Pro” label therefore does not designate a version that automatically retains and improves every H3 capability. It instead identifies the quality-focused model in Pruna’s video lineup, built around a narrower interface and two generation settings.
Speed mode is enabled by default. It prioritizes fast generation and is intended for repeated tests, animated storyboards, or alternative versions to compare. Quality mode takes more time and costs roughly twice as much, with the promise of higher visual fidelity.
The documentation does not yet quantify that difference. It provides no generation times for Quality mode, no measurements of additional consistency, and no detailed comparison between the two outputs. For now, the choice is therefore based on Pruna’s stated positioning rather than reproducible results.
First-frame control can establish the initial appearance of a character, product, or setting. The prompt then describes the expected motion, actions, camera movement, lighting, or style of the sequence.
Adding a final image provides a second fixed point. The model must create the motion needed to connect the initial state to that ending. This feature can support transformations, transitions between two compositions, or generations that follow an existing storyboard more closely.
It does not, however, provide precise control over everything that happens between the two endpoints. A character’s path, the speed of a movement, or an object’s transformation are still interpreted by the model. Two compatible images reduce ambiguity, but they do not replace frame-by-frame animation.
A text prompt remains mandatory even when a starting image is provided. It can specify actions, atmosphere, camera movement, and expected sound events. Requests that conflict with the reference composition may nevertheless produce distortions or less stable continuity.
Seven aspect ratios are supported: 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, and 1:1. When a reference image is supplied, its proportions determine the canvas directly, making the aspect-ratio parameter secondary.
Users can also set a seed to make results easier to reproduce. This is mainly useful for comparing different settings from a similar starting point. It does not necessarily guarantee that a video will remain strictly identical after updates to the model, execution engine, or provider infrastructure.
Pruna adds another setting independent of Speed and Quality modes. The `promptupsampler` automatically expands the prompt before generation. Three levels are available: Off, Turbo, and Max, with Turbo enabled by default.
This step can enrich a short request by adding details about the camera, lighting, setting, or movement. It can also shift the result away from the original intention if the rewrite introduces choices that were never requested. The Off setting therefore remains useful for teams seeking more literal control over their prompts.
Pruna says the prompt-expansion level does not affect pricing. It may still affect total latency because the time spent rewriting the text is added to the video-generation process. The reported 1.99- and 4.3-second figures appear to refer to model inference rather than the entire journey from submitting a request to receiving the file.
The delay experienced by users also includes queue time, prompt processing, encoding, storage, and file transfer. Pruna has not yet published end-to-end measurements, median results across a large number of requests, or latency figures during peak demand.
Performance is not detailed across every configuration either. The two highlighted figures apply to a five-second video in Speed mode. No times are provided for 10- or 15-second clips, Quality mode, or requests using both starting and ending images.
The service generates a soundtrack with every video. MiniMax H3 was designed to produce visuals and sound together, which can improve the relationship between an action and its