H3 combines generation, multimodal references, and video editing
The MiniMax H3 model combines generation, local editing, and the integration of twelve visual or audio references to stabilize shots.
MiniMax H3 combines text-to-video generation, image animation, first-and-last-frame control, and the use of visual, audio, and video references within a single context. The model produces sequences up to 2K, with natively generated stereo sound and a duration of up to 15 seconds.
The reference mode accepts up to nine images, three video clips, and three audio files, within an overall limit of twelve elements. These sources can be used to maintain a character, replicate a camera movement, transfer a voice, follow an editing rhythm, or maintain a visual direction throughout the generation. Prompts can be up to 7,000 characters, allowing for the description of multiple shots or relationships between references in a single request.
H3 also supports localized editing of existing sequences. It can replace a product, modify text, change a line of dialogue, adjust lighting, or add an element without intentionally rebuilding the entire shot. MiniMax particularly emphasizes the rendering of typography, interfaces, subtitles, and animated graphic elements.
This approach is based on an architecture that jointly processes text, image, video, and audio. To produce 2K outputs, the model regenerates its initial version while preserving the original context, rather than solely applying a separate upscaling step. MiniMax claims this process helps recover fine details and better maintain text.
H3 is accessible via API, as well as in ComfyUI through a partner integration. MiniMax also plans to make the model weights accessible soon, subject to applicable legal constraints.