Voice Design accelerates the casting of virtual agents and characters

Gradium launches Voice Design, a tool that generates synthetic voices from a description in just a few seconds, featuring five languages and regional accents.

A composed British receptionist, an older American narrator, or a warm-voiced Québécois advisor. With Voice Design, Gradium lets users describe the performer they need before that voice even exists.

The tool takes a written prompt specifying characteristics such as accent, apparent age, gender, pace, pitch, energy, or register. It then produces up to five voice options within seconds. Users can listen to them, generate another batch, or save the one that best fits their project.

Unlike a traditional catalog, the process no longer involves browsing through a list of prerecorded voices. It begins with a creative brief written in natural language. Gradium’s examples include a pirate from Bristol, an Irish customer service representative, and an older American narrator.

Each request generates new results, even when the description remains unchanged. Voice Design does not retrieve an existing voice from a library. The model samples a new option instead. Consistency begins only after a voice has been selected. Once saved, its `voiceid` remains stable for as long as the voice is retained.

This distinction matters for productions designed to continue over time. A team can explore several interpretations at the beginning of a project, then lock in a voice so that later episodes, conversations, or messages retain the same sonic identity. Without saving it, a new generation based on the same prompt will not reproduce the exact previous result.

Gradium directly contrasts this method with voice cloning. Cloning begins with a recording and attempts to reproduce the voice of a specific person. Voice Design starts only from a written description and creates a synthetic voice that, in principle, does not belong to any recorded performer.

The two features therefore address different needs. Cloning is useful when a project must recreate a person who has already been selected, with their consent and the necessary rights. Voice Design applies when the role, accent, or personality has been defined, but the speaker has not.

The lack of an initial recording simplifies some contractual requirements. Gradium says a voice created with its tool does not require a voice actor license, renewed consent, or voice-specific royalties. The customer retains the rights to the audio files they generate, to the extent that those outputs are eligible for protection under applicable law.

A fully synthetic voice does not, however, give users the right to imitate a real person without permission. Gradium’s Terms of Service prohibit uses that infringe publicity, personality, privacy, or similar rights, as well as impersonating a person or company without explicit authorization.

A prompt that simply requests a deep, slow, Québécois voice therefore raises different issues from one intended to recreate an identifiable public figure. Responsibility for both the submitted description and the final use remains with the customer.

Ownership of a generated file does not guarantee complete exclusivity over the voice either. Gradium states that outputs from its services may not be unique and that other customers may receive identical or similar results. A `voiceid` makes a voice stable within the service, but it does not automatically turn that voice into an exclusively reserved sonic identity.

Voice Design currently supports English, French, German, Spanish, and Portuguese. Prompts can specify regional accents within those languages. Gradium recommends naming a precise region, such as “Colombian” instead of “Latin American,” or “Québécois” instead of “French Canadian.”

This level of detail reduces ambiguity, but it does not guarantee that every listener will identify the accent in the same way. An accent contains geographic, social, and generational variations that are difficult to capture in a few words. The result should therefore be reviewed by people familiar with the region, particularly when it is meant to represent a community or local brand.

Saved voices then enter Gradium’s regular workflow. Their identifiers can be sent to the same streaming text-to-speech endpoint used for catalog voices and clones. They support the same output formats and do not require a separate production pipeline.

This continuity is aimed at voice agents, games, audiobooks, virtual characters, automated phone systems, and customer service applications. A company could prepare several identities for different markets, while a studio could create supporting characters without immediately organizing a recording session.

Prompt-based design also simplifies early testing. A team can compare a calm interpretation with a more energetic one, adjust the apparent age, or refine the accent before integrating the voice into an application. Listening remains essential, however. A well-written prompt cannot by itself verify the pronunciation of a name, emotional consistency, or the quality of a long conversation.

Gradium presents Voice Design as a free feature available through its API and Studio. That wording refers to access to the tool, but it does not mean that every generated file can be used commercially without a subscription.

The company’s pricing page reserves commercial use for paid plans. The free tier supports five custom voices and includes 45,000 credits, which the company estimates is equivalent to around one hour of speech generation, but it is limited to testing and internal, noncommercial use.

The first paid plan, XS, costs $13 per month. The following plans are listed at $43, $340, and $1,615 per month, with additional credits and more simultaneous connections. Standard paid subscriptions can retain up to 1,000 custom voices, compared with five on the free tier. The Enterprise plan lists no limit.

There is therefore no separate royalty charged for each synthetic performer, but the service itself remains subject to subscription fees and credit usage. The promise of commercial use without voice licensing should be understood within that framework.

To accompany the launch,