Scene extension and 4K export arrive in Gemini Omni 1.1 Flash

Gemini Omni 1.1 Flash extends videos to 40 seconds, connects two images, speeds up drafts, and exports in 4K.

Extending a sequence, setting its opening and closing shots, or preparing several lightweight drafts before a final render: Gemini Omni 1.1 Flash adds controls intended to bring video generation closer to a production workflow. The model is now generally available through the Gemini API, following the introduction of Gemini Omni Flash at Google I/O and its public preview release in late June.

The main addition is scene extension. An existing video can now be continued in ten-second increments, up to a cumulative length of 40 seconds. Gemini Omni 1.1 Flash analyzes the previous segment’s final ten seconds to preserve characters, movement, surroundings, and audio continuity. The initial version referenced only the last second.

Each extension receives a new instruction. Creators can continue a conversation, move the camera, gradually change the setting, or take the story in a different direction. The model may slightly alter the final frames of the original segment to make the transition less noticeable.

The 40-second limit does not represent a single generation. Clips are created or extended in segments lasting three to ten seconds. These sections are connected through the Interactions API, which retains the generation’s state using the previous interaction ID. An application can create an opening scene, request a continuation, and submit further instructions without resending the entire project at every step.

Extensions can only be added to the end of a clip. The model cannot yet place a new scene before a video or insert footage into the middle of an existing sequence. When an outside file is uploaded, its duration cannot exceed ten seconds. This restriction does not apply in the same way to videos generated and extended within a single conversation.

Dialogue handling also remains limited. A video created by the model can receive additional spoken lines when extended through multiple interactions. However, users cannot upload footage of someone speaking and then make that person deliver a new line in the generated continuation. Direct voice editing is not supported.

Gemini Omni 1.1 Flash also adds interpolation between a first and last image. The developer supplies both ends of the shot and describes the intended transition. The model then produces the movement between them, such as an orbit around a subject, a push-in, a tracking shot, or a loop returning to the opening image.

This feature is aimed at productions where the visual destination needs to be established in advance. A creator can define the beginning and end of a camera movement without leaving the result entirely to the interpretation of a text instruction. It does not guarantee a mechanically exact transition, however. The intermediate images are still generated and require review before use.

The model also accepts video references. Up to three clips, each lasting no more than three seconds, can be used to reproduce the appearance of a character, object, or movement. Google demonstrates the feature with several animals assigned dance routines taken from separate reference videos. Audio from those references is ignored, and the documentation advises against asking the model to perform complex reasoning across multiple clips.

Developers can now choose between four output resolutions: 360p, 720p, 1080p, and 4K. The default remains 720p. The 1080p and 4K versions are produced by upscaling the output rather than generating it natively at those resolutions. That distinction matters for fine details: a larger export does not necessarily recreate information missing from the initial result.

The 360p mode is intended for quickly exploring several directions. Google claims it generates video up to 60% faster than 720p, based on its own system throughput, at one-third of the cost. A team can produce several drafts, compare them, and reserve high-resolution rendering for the selected version. Actual gains will depend on service load, requested duration, and the number of attempts required.

Standard API pricing is calculated in tokens. Google charges $1.50 per million input tokens across all modalities, $9 for text output, and $17.50 for video output. One second of 720p video represents 5,792 tokens, for an effective cost of approximately $0.10. A ten-second sequence therefore costs close to $1 for video output alone, before inputs and any additional generations. The API offers no free tier for this model.

Video and audio are generated in the same response. Instructions can specify an atmosphere, music, dialogue, sound effect, or the exact moment when an event should occur. A prompt might request a scene change after three seconds or the start of a chorus at the five-second mark. These directions guide the model but do not offer the temporal precision of editing software.

The conversational editing introduced with Omni Flash is still available. An application can submit a video, ask the model to remove an object, alter the lighting, or change the visual style, then continue making revisions within the same history. Google recommends short, focused instructions accompanied by an explicit request to keep everything else unchanged. This reduces unintended alterations without eliminating them.

The description of Omni as a natively multimodal model requires some qualification in light of the current API. Gemini Omni is designed to process text, images, video, and audio, but the documentation does not yet support uploading a separate audio track as a reference. Sound from video references is also discarded. Audio generation is therefore guided primarily by text and the context of an earlier generation.

Some functions differ by region. In the European Economic Area, Switzerland, and the United Kingdom, the API currently cannot edit or extend an uploaded video. Users can still continue footage generated by Gemini Omni within the service. Other restrictions apply to images of minors and certain recognizable people. English is the only language Google says it