Mercury Voice brings reasoning below 320 ms for voice agents
Inception is opening Mercury Voice to enterprise customers, a diffusion LLM built for voice agents. The model reaches a 320 ms median time to first answer token in the company's tests, with three reasoning levels, a 128K context window, and an OpenAI-compatible API.
320 ms before the response begins
Voice conversations leave little room for silence. Inception estimates that an agent has roughly 500 ms after a caller stops speaking to begin responding before the pause starts to feel unnatural.
Mercury Voice targets that window directly. On a set of production voice-agent prompts, Inception measures a 320 ms median time to first answer token after reasoning, with a p95 of 750 ms.
The metric, Time to First Answer Token or TTFAT, does not simply measure when generation begins. It waits for the first token that will actually become part of the spoken response after reasoning is complete.
Under this protocol, Inception reports 5.9x lower median latency than GPT-6 Luna without reasoning. These results come from the company's own evaluation. A diffusion LLM built for conversation
Mercury Voice belongs to Inception's family of diffusion LLMs. Instead of relying on strictly sequential autoregressive generation, Mercury works on multiple tokens in parallel to reduce inference time.
The model retains capabilities expected from an agent, including reasoning, tool calls, and long system-prompt following. Three reasoning effort settings are available: Low, Medium, and High.
Its context window reaches 128K tokens, with up to 50K output tokens.
Two weeks earlier, Mercury 2.5 provided the first preview of the voice-specific model. Inception was already targeting workloads where a few hundred milliseconds directly affect how an interaction feels. Quality measured through agent tasks
The comparison is not limited to speed. Inception aggregates several conversational and agentic evaluations, including τ³-bench Telecom, Retail, and Airline, IFBench, and BFCL v4.
According to its results, Mercury Voice scores above Gemma 4 31B, GPT-6 Luna, GLM-5.3-Flash, and Qwen3.5-397B on the composite.
Across the three τ³-bench categories, Inception also reports higher average quality than almost every model it tested while running twice as fast as the next fastest model.
These comparisons were produced by Inception and are not independent evaluations. Early deployments in phone calls and drive-throughs
Audivi AI is using Mercury Voice for automated drive-through ordering, including order changes and sales interactions during live conversations.
Altur uses it for autonomous phone agents serving financial institutions, where calls can involve payment-plan negotiations and objection handling.
OpenCall reports median model response latency close to 170 ms on its own production workload. That figure reflects its specific environment and should not be conflated with the 320 ms measured on Inception's cross-model voice prompt set. Under one cent per minute, according to Inception
Standard pricing is $0.40 per million input tokens and $1.50 per million output tokens.
A 50% launch discount temporarily brings those rates down to $0.20 and $0.75. For what Inception describes as a typical voice-agent profile, the company estimates a cost of approximately $0.009 per minute of conversation.
It compares that figure with an estimated $0.045 per minute for GPT-4.1 under the same profile.
Mercury Voice is available to enterprise customers through the Inception API using an OpenAI-compatible endpoint. It can therefore occupy the LLM layer in voice-agent stacks built with LiveKit, Pipecat, Vapi, Retell, or custom infrastructure. Access currently goes through Inception's enterprise sales channel.