ElevenCreative brings together generation and editing in Studio 4.0
ElevenLabs transforms Studio into a complete video production space within ElevenCreative. Image, footage, voice, music, and sound effects generation now join multitrack editing, subtitles, and collaboration.
A shot is missing in the middle of the edit. It can now be described, generated, and placed on the timeline without opening another interface. With Studio 4.0, ElevenLabs brings asset creation, editing, and team feedback into the same video project.
The update applies to Studio, the editor built into ElevenCreative. The software now combines the image and video models available on the platform with ElevenLabs’ voices, music, and sound effects. Generated assets arrive directly in the project where they are intended to be used.
The change is less about adding another isolated generator than removing the transfers between separate workspaces. A user can start with a script, produce the necessary shots, add a voiceover, create a soundtrack, correct the captions, and export the result without repeatedly downloading and re-uploading each element.
This consolidation continues a shift that has been underway for several months. ElevenLabs previously added several image and video models to its platform, then introduced Flows for building reusable production workflows. Studio Agent also joined the editor to prepare an initial cut through conversation.
Studio 4.0 does not introduce this co-editor from scratch. Studio Agent was already available, but it now operates inside a video environment rebuilt for projects containing multiple scenes, tracks, and asset types.
A project can begin with existing media or an empty timeline. Studio supports several common video formats, then automatically selects a suitable layout. Guided workflows can also help users create faceless videos, add voiceovers, generate soundtracks, produce captions, or dub footage.
Generation now takes place inside the edit. A description can be used to create a video, still image, narration, music track, or sound effect. Studio can also retrieve previous creations stored in an ElevenCreative workspace.
ElevenLabs states that every ElevenCreative model can be used from the editor. Their capabilities still vary. Available durations, formats, resolutions, reference inputs, and audio options depend on the selected model. Some models are also unavailable in certain countries.
Studio Agent does not automatically use the full catalog when selecting a model on its own. The documentation states that it chooses from five predefined image models and five predefined video models. Before generation begins, the user can review this choice and replace it.
This creates two separate levels of access. The editor provides the wider ElevenCreative catalog, while the co-editor relies on a smaller selection when deciding which model best matches a request.
An instruction such as “create a 30-second teaser from this footage” may prompt Studio Agent to ask about length, tone, structure, or transitions. It can then prepare an initial cut, arrange the footage, choose a voice, generate the narration, and add selected audio elements.
Two modes govern its actions. In Plan mode, the agent analyzes media, transcribes dialogue, searches for assets, and describes the operations it intends to perform without changing the timeline. Create mode allows it to insert clips, generate content, adjust audio, and apply text overlays.
This separation addresses a common issue with assistants capable of acting inside software. An imprecise instruction can trigger a series of changes that are difficult to control. Plan mode gives users an opportunity to review the proposed approach before it is applied to the project.
The timeline remains manually accessible at any time. A user can interrupt the agent, move a shot, shorten a sequence, change a voice, or revise the overall pacing. Studio Agent can then continue from the updated version.
The workflow therefore alternates between delegation and direct editing. Users are not confined to a conversation in which every correction must be written out. They can intervene with conventional editing tools, then return control to the co-editor for a broader task.
Studio Agent also analyzes videos frame by frame to build a temporal map of their content. A request can target a specific event, such as adding a sound when a product enters the frame or displaying text when a gesture occurs.
This analysis does not mean every decision will be accurate. Detecting an object, shot change, or action can remain ambiguous. ElevenLabs has not published an evaluation measuring placement accuracy, the number of corrections required, or the consistency of results across long projects.
The timeline has been redesigned to support frame-level cuts, clip snapping, and the movement of multiple clips while preserving synchronization. Narration, video, music, and sound effects appear on separate tracks.
Waveforms make it easier to align audio levels and events. Clips can be trimmed, split, duplicated, or repositioned. Fades, volume, and various narration settings remain accessible from a sidebar that adapts to the selected asset.
Captions also become timeline elements. A line can be split, merged with the next one, or corrected without restarting the entire transcription. Both the text and its duration can be edited directly before export.
Studio provides visual templates for caption fonts, colors, and placement. The captions are then burned into the exported video. For this feature, the documentation does not mention the separate creation of professional caption files intended for an external broadcast workflow.
The ability to “refine” a shot without starting over should also be understood as continuity within the project. A new asset can replace the previous one in the same position without rebuilding the surrounding edit. This does not guarantee that a model will preserve every detail unaffected by the requested change.
A new generation remains a new operation, with a result that may vary. Its duration, cost, and available controls depend on the model being used. Visual features do not necessarily receive the same free regeneration allowances available for certain narration tasks in