P-Video-2 claims it can generate a video faster than it plays
Pruna AI launches P-Video-2, a model capable of producing up to 20 seconds of 1080p video from text, an image, or an audio file, featuring a lower-cost draft mode.
Start generating a 20-second sequence and receive it before it would have finished playing once. That is the promise behind P-Video-2, Pruna AI’s new model for short-form video generation.
Introduced on September 10, 2026, the system accepts a text prompt, a starting image, or an audio file. It produces sequences up to 20 seconds long in 720p or 1080p at 24 or 48 frames per second.
Pruna reports generation times starting at approximately 0.41 seconds per second of video in Draft mode and 0.91 seconds in Standard mode. At those rates, a 10-second clip would theoretically take about 4.1 seconds in Draft or 9.1 seconds at standard quality. A 20-second sequence could take 8.2 or 18.2 seconds.
These figures are starting points published by the company, not guaranteed turnaround times for every request. Pruna has not yet specified the output settings, accelerator, batch size, or testing method behind the measurements. Time spent in a queue, file transfers, and downloading the result may also increase the overall delay.
The description of P-Video-2 as the “fastest video generation model” should therefore be treated as a marketing claim. Pruna says it worked with Datapoint, Rapidata, and DesignArena to validate its evaluations, but the complete results have not been released. No public table currently compares P-Video-2 with competing services using identical prompts, settings, and hardware conditions.
Pruna’s own model benchmark page does not yet include P-Video-2. It also warns that response times may differ considerably between providers and recommends testing each endpoint on the workloads for which it will actually be used.
Standard mode targets final output. Draft mode reduces some of the quality to return a preview more quickly. A studio can test the framing, motion, duration, or overall direction of a scene before running the request again at the higher quality level.
This separation avoids paying the maximum rate for every attempt. It does not guarantee that a draft and its final render will be perfectly identical. A seed can limit variations, but switching quality levels changes the generation process and may affect certain movements or details.
The official P-Video-2 documentation sets the duration between one and 20 seconds. That field can be left empty, allowing the model to select a length based on the prompt instead of always imposing a five- or 10-second sequence.
The option adds flexibility but makes the final price less predictable. Pruna bills customers for the duration of the video returned, measured in whole seconds. When the model selects the length, the exact charge is therefore known only after generation is complete.
At 720p, Pruna’s direct pricing is $0.025 per second at standard quality and $0.015 in Draft mode. A 20-second video costs $0.50 or $0.30, respectively.
Moving to 1080p doubles those amounts. The price reaches $0.05 per second in Standard mode and $0.03 in Draft, bringing a maximum-length clip to $1 or $0.60. The public pricing table does not list a separate surcharge for choosing 24 or 48 frames per second.
The pricing documentation contains a contradiction. The description of the `draft` parameter says it is billed at half the standard rate. The four listed prices show that Draft actually costs 60% of the standard price, representing a 40% discount. Until the wording is corrected, customers will need to rely on the numerical table and the rate displayed in their dashboard.
P-Video-2 is also more expensive than its predecessor. P-Video was priced at $0.02 per second in 720p and $0.04 in 1080p, with Draft rates of $0.005 and $0.01. Standard prices have increased by 25%, while Draft prices have tripled.
In return, the advertised maximum duration increases from 15 to 20 seconds, and standard generation becomes considerably faster according to Pruna’s figures. The first generation took about 10 seconds to produce five seconds of 720p video, or close to two seconds of processing for each second of output. P-Video-2 now claims a starting rate of 0.91 seconds.
The three input types cover different uses. A text prompt creates the scene from a description alone. An image establishes the initial appearance of a character, product, or setting. An audio file can guide the pacing, speech, or movement.
Supported image formats include JPEG, PNG, and WebP. When an image is supplied, the aspect-ratio setting is ignored and its dimensions determine the general shape of the result. Without an image, seven formats are available, including 16:9, 9:16, 1:1, 4:3, and 3:2.
Users can also provide a second image representing the intended final frame. The model must then generate a transition between the starting point and that reference. This control could be used to prepare a cut, show a product transformation, or create a shot whose final composition has already been decided.
It is not a strict geometric constraint. The system interprets both references and invents the intermediate steps. A complex final image that differs sharply from the first or conflicts with the requested motion may result in an abrupt transition or visible distortions.
For audio input, the API accepts FLAC, MP3, and WAV files. When one is supplied, its length overrides the value entered in the `duration` field. Billing also follows the length of the audio. The documentation does not yet explain what happens when a file exceeds the general 20-second limit, whether it is rejected, shortened, or processed automatically.
P-Video-2 can also create a soundtrack from the prompt. The listing prepared by Segmind mentions the generation of dialogue, music, and sound effects alongside the video, with lip synchronization. Pruna’s `saveaudio` parameter, enabled by default, can remove that audio track from the delivered file.
A user-supplied audio file therefore does more than add a soundtrack after the visuals have been created. It conditions the generation in an attempt to align gestures, mouth movements, or editing with its content. The