Inception made Mercury Voice available to enterprise customers on September 29. Its headline result is a company-reported 320-millisecond median time to first answer token, with a 750-millisecond p95, at low reasoning effort on production voice prompts. Those figures describe the language-model stage of a call. They do not establish how long a caller waits before hearing a useful spoken response. Source: Inception’s launch announcement.
That distinction should shape the buying decision. A faster model is useful when the model is where the delay lives. If an order lookup takes two seconds, changing the model will leave that lookup in the critical path.
This is a review of published evidence and an evaluation plan, not a hands-on Mercury Voice test. We have not measured the model or placed calls through it.
What launched, and how you get it
Mercury Voice is a diffusion language model for voice agents, with reasoning and tool use. Inception’s model page lists a 128K-token context. The launch describes an API compatible with the OpenAI request format that fits into the language-model slot of a voice stack. Launch details.
The original X announcement also directs enterprise customers to contact the company for access. That is the supported starting point. A compatible request format should not be treated as proof that your existing provider account, rate limits, tool schemas, or every integration will work unchanged.
Ask for the exact model identifier, supported request fields, regional routing, and access limits before writing a migration. Keep those answers with the model version you evaluate. An integration that silently falls back to your previous model would tell you little about Mercury Voice.
Where the caller’s waiting time goes
A voice agent has to decide that the user finished speaking, obtain the relevant text, produce an answer, turn that answer into audio, and deliver it. A tool call can add another wait. Some stages overlap, which is why adding unrelated benchmark numbers often produces a misleading total.
LiveKit’s observability documentation distinguishes per-component metrics from per-turn measurements. It exposes end-to-end latency through conversation metrics, alongside transcription, turn-completion, language-model and speech-synthesis measurements. Its component definitions also note that end-of-utterance delay already includes transcription delay. Counting both again would inflate a pipeline estimate.
For a practical evaluation, record two clocks. The first measures the model’s answer start. The second begins when the caller stops speaking and ends when the caller hears the response. Keep the audio recording so a reviewer can distinguish a meaningful answer from a quick filler phrase.
Then label each turn by what happened: direct answer, tool lookup, correction, interruption, or escalation. An average across those categories can conceal the exact interaction that makes a support call frustrating.
Measure successful tasks alongside speed
Consider a synthetic order-change scenario: the caller removes one item, adds another, then corrects the delivery address. A useful test records whether the final order and address match that request, whether the agent confirms material changes, and whether a repeated tool response creates a duplicate order.
Use a separate synthetic account-recovery scenario. The agent should request the required verification before revealing account information. Score an unauthorized disclosure as a failure even if it arrives quickly. These are suggested test cases, not observed Mercury Voice results.
Compare the same prompts, tool responses, and audio conditions across models. Report both median and slow-tail waiting time, task completion, incorrect tool actions, and human handoffs. Preserve failed turns. If you change the prompt between runs, document the change rather than attributing the entire improvement to the model.
Inception’s performance claims remain vendor evidence. A useful internal pilot would test the work your callers actually need done, rather than assume a general benchmark result predicts that work.
The token bill is only one part of the call
The launch lists standard rates of $0.40 per million input tokens and $1.50 per million output tokens, with introductory rates of $0.20 and $0.75. Confirm that the promotion still applies when access is granted. Inception pricing.
There is a documentation discrepancy: the Mercury Voice section of the model page displays different prices and an under-170 ms time-to-first-token claim. The dated launch uses time to first answer token. These pages do not establish that the measurements are comparable. Use the dated announcement for the calculation below, and obtain an account-specific quote before budgeting a deployment.
Here is a bounded calculation using those published rates: one million input tokens plus 200,000 output tokens costs $0.35 at the introductory rates, or $0.70 at the listed standard rates. That is arithmetic for an assumed token workload, not a measured call cost. Extra model attempts would increase it.
Budget separately for transcription, speech generation, telephony, hosting, and any tools called during the conversation. Track cost per successfully completed task as well as cost per minute. Ending a cheap call with an incorrect change can make the whole workflow expensive to repair.
Our n8n Agents explainer makes a related point about bounding agent actions: choosing a tool and completing the requested workflow are separate responsibilities.
A sensible first pilot
Start with one reversible workflow using synthetic records. Keep the previous model available as a documented fallback, put an operator in charge of escalation, and decide the acceptance criteria before collecting results. Include noisy audio, a caller who pauses mid-sentence, a correction during the agent’s reply, and a delayed tool response.
Retain the model identifier, reasoning setting, prompts, tool traces, token usage, and audio for each test. Review failures before widening the pilot. Mercury Voice deserves that evaluation because its new availability makes the model a concrete option; the decision to deploy should rest on your complete spoken workflow.
The Kingy Brief
Get the next Kingy Brief.
Source-checked AI changes, original tests and one practical thing to try.
Free · Choose your subjects · Double opt-in · Unsubscribe anytime
