ElevenLabs' MCP connects assistants to over 50 creative models
ElevenLabs' Model Context Protocol connects AI assistants to 50+ creative models, enabling seamless integration of advanced voice and audio generation.
Create an image, animate it, add a voice, music, and sound effects, all without leaving the conversation with your assistant. The ElevenLabs MCP now brings these operations together through a single connection.
The connector’s expanded scope covers speech synthesis, transcription, dubbing, music, sound effects, image generation and editing, video, and lip-syncing. Results appear directly in ChatGPT, Claude, Cursor, Grok Bot, or Hermes, then are saved to the ElevenCreative workspace associated with the account.
The first version, released a few weeks earlier, focused on ElevenAgents. It allowed users to review the performance of voice and text agents, create new ones, change their configurations, or estimate language model costs. The same connector now provides access to creative features, with no additional installation required if the integration was already active.
MCP, short for Model Context Protocol, provides a common interface between an assistant and external services. ElevenLabs models are not built directly into ChatGPT or Claude. The assistant translates a request into a tool call, the hosted server executes the operation on the relevant infrastructure, and the resulting file is returned to the conversation.
The connection uses OAuth. The user selects ElevenLabs from the assistant’s connector or plugin directory, signs in to their account, and authorizes access to the workspace. No API key needs to be pasted into the conversation, and no local server is required.
The setup process varies slightly by client. ChatGPT, Claude, Cursor, and Grok Bot support installation from their respective directories. Claude Code and Hermes require users to add the server address from a terminal before authenticating. The official MCP page provides the commands and links for each environment.
For speech synthesis, a request can select an available voice from the library, specify the text, and describe how it should be delivered. The audio file is returned to the conversation and remains available in ElevenCreative. Private or shared voices remain subject to workspace permissions.
Scribe handles the reverse process. A recording sent to the assistant can be converted into text with speaker identification and timestamps. ElevenLabs claims support for 99 languages in its MCP presentation. The transcript can then be used as the basis for captions, a script, or a dub.
Dubbing runs through the same connection. The user provides content and requests another language. The service attempts to preserve the original speaker’s voice, tone, and delivery, then returns the finished track to the conversation.
Music and sound effects complete the audio workflow. An instruction can request an instrumental track or a song with lyrics, define its style, mood, or duration, and generate several effects for a scene. Each operation remains separate, although the assistant can chain them together from a shared brief.
The visual features come from ElevenCreative Flows. The Image & Video documentation lists GPT Image 2, Nano Banana 2, Krea 2, Seedream, Sora 2, Veo 3.1, Kling, and several Runway models, among others. The catalog covers image generation and editing, text-to-video creation, image animation, video-to-video transformation, and lip-syncing.
Available features vary by model and region. Several services from ByteDance or Kling are unavailable in the United States. Some features also require specific visual references, enforce fixed durations, or limit output formats. The claim of more than 50 models therefore does not mean every account has access to the same catalog.
A single request can combine several steps: write a script, generate the voiceover, compose a soundtrack, prepare sound effects, create the visuals, and then produce the video. The official announcement presents this workflow as a way to move from one brief to a finished piece.
The result is better understood as a first draft. The announcement itself directs users to ElevenCreative Studio to revise the timeline, adjust the narration, edit the audio tracks, and prepare the export. Bringing the tools together reduces the need to move between services, but it does not eliminate editing or editorial review.
Every creation is added to the ElevenCreative workspace. It can then be opened in Studio or incorporated into a more detailed Flow. The conversation serves as the command interface, while the platform stores the files and provides the tools needed for revisions.
Every step available in Flows is intended to be accessible through the MCP. This environment lets users build multimodal workflows on a canvas, then rerun only the part that needs to change. A new voice can, for example, replace the previous one without regenerating images that do not depend on the audio track.
ElevenCreative Flows remains in Alpha, while image and video generation are labeled as beta features. Capabilities, models, and access conditions may still change. API execution of Flows for large-scale automated production is also planned for a future release.
The connector retains its agent management capabilities. From the same assistant, an authorized user can review conversations, compare configurations, edit a prompt, replace a voice, create an agent, or delete one. Creative access does not replace the original MCP. It expands the range of actions available through it.
This expansion makes permissions more important. Administrators choose which tools the connection is allowed to call. In Enterprise workspaces, visual models are disabled by default and must be authorized individually through the model approval settings.
OAuth authentication prevents an API key from being exposed, but it does not stop an authorized operation from changing a configuration or consuming credits. The connector’s access level should match the user’s role and the degree of control required over the workspace.
There is no separate price listed for installing the MCP. Generations are billed according to the selected model and