fal turns MiniMax H3 into a faster video model

H3 Max is a fal-post-trained version of MiniMax H3, available for text-to-video and image-to-video generation at 480p or 768p with synchronized audio.

H3 Max is neither a standard MiniMax model merely distributed by fal nor an entirely independent creation. It is a new version developed by fal Research from MiniMax H3, whose weights were released by MiniMax. fal used this base model, applied additional post-training, and adapted it to its own inference engine.

This distinction explains both the “H3 Max by fal” branding and its placement under the `minimax/h3-max` endpoints on the platform. MiniMax remains responsible for the original architecture and base model, while fal claims responsibility for the additional training data, alignment work, quality improvements, and infrastructure used to serve this version.

The post-training focused on instruction following and visual quality. fal says it introduced a significant amount of new data and devoted part of its training compute to reinforcement-learning tasks with verifiable outcomes. The company has not disclosed the composition or exact size of the dataset, nor has it provided a detailed account of the procedures involved.

The model was developed alongside its inference engine. fal says it retained only optimizations that did not reduce quality in its internal evaluations. H3 Max was trained and deployed on NVIDIA GB200 NVL72 systems, but it remains a remote service: users do not need specific hardware to access it.

H3 Max’s main claim concerns speed. According to fal, a five-second 768p video can be processed in less than three seconds. One example in the documentation reports approximately 2.77 seconds of inference time. This figure measures server-side processing rather than the complete delay experienced by the user, which may also include queue time, file transfers, and result retrieval.

fal estimates that this performance represents roughly 35 times the throughput of the official MiniMax H3 endpoint and an average speed 15 times higher than the models of comparable quality included in its tests. The gain does not remain identical at every duration: the product page states that a 15-second sequence may take approximately 15 seconds to process.

H3 Max supports generation from either text or an image. Users can provide an opening image and optionally a closing image to guide the transition between two frames. Clips can run from five to 15 seconds and are generated at 480p or 768p, at 24 frames per second, with synchronized audio produced in the same pass.

For text-to-video generation, the available aspect ratios include 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16. For image-to-video, the output follows the aspect ratio of the input image. H3 Max does not provide the 2K output available through the standard MiniMax H3 workflow. Its priority is speed and production scenarios requiring a high volume of generations.

To evaluate quality, fal compared the model against 12 other video systems. Human evaluators selected the stronger result across three criteria: overall preference, prompt understanding, and aesthetics. The comparisons were then aggregated into Bayesian Elo ratings with 95% confidence intervals.

H3 Max ranked first across all three categories in fal’s internal evaluation. These results should be treated as measurements produced by the company behind the model, even though human comparison is relevant in a field where automated metrics capture only part of the viewing experience.

Public leaderboards provide a second reference point. Design Arena gave H3 Max an Elo score of 1,341 for image-to-video generation, ahead of MiniMax H3 at 1,333. Artificial Analysis listed it under the internal name “MiniMax H3 Turbo 768p” with audio, assigning it an Elo score of 1,201 with a confidence interval of plus or minus 11 across 2,177 evaluations.

These rankings may change as new votes and models are added. fal also acknowledged that an initial launch graphic displayed incorrect Artificial Analysis rankings for several competing models. The company later corrected the positions of Gemini Omni Flash, Kling 3 Pro, Wan 3.0, Seedance 2.5, and FLUX.3. The correction makes the continuously updated public leaderboards more reliable references than the original graphic.

H3 Max is available through fal’s Playground, fal Agent, and API. The documentation states that its outputs can be used commercially. fal also offers a free tool providing five daily five-second 768p generations without registration, plus five additional generations through the Sandbox after signing in.

Until September 1, 2026, launch pricing is set at $0.025 per second for 480p and $0.04 per second for 768p. A five-second video therefore costs $0.125 or $0.20, respectively. After the promotional period, those rates are expected to increase to $0.05 and $0.08 per second.

MiniMax H3 remains an open-weight model that can be downloaded, modified, and run outside fal’s services. H3 Max has not received an equivalent release: no download for its weights appears in the announcement or documentation. The release reflects an expansion of fal’s role, moving beyond hosting third-party models to developing its own versions from open-weight foundations.