← AI Switchboard
AI Switchboardby Waggle
MODELS · September 30, 2026
Sep 29

Inception makes Mercury Voice generally available at half price, claiming a first word in under 320 milliseconds

A diffusion language model built for voice agents, at $0.40 and $1.50 per million tokens, discounted to $0.20 and $0.75 at launch.

Inception made Mercury Voice generally available to enterprise customers on 29 September, two weeks after previewing it. Inception's Mercury models are diffusion language models, a different design from the usual word-by-word generators, and this one is aimed at voice agents such as phone assistants.

The list price is $0.40 per million input tokens and $1.50 per million output tokens. At launch it is 50% off, at $0.20 and $0.75. The post gives no end date for the discount.

The performance pitch is speed. Inception says “Mercury Voice returns its first answer token in under 320 milliseconds (median), while also beating models like GPT-6 Luna and Gemma 4 31B.” That is the company's own measurement; no outside test has been published.

Why it is here: latency is what a caller notices first, and voice is where agent pricing is being cut hardest. The price war covered here on 24 September, when Qwen cut voice prices by up to 95% and Google shipped designed voices, now has a new entrant with a different architecture. For businesses building phone agents, a sub-second reply at these prices lowers the cost of trying one.

  • Confirmed Generally available for enterprise customers on 29 September. Inception
  • Confirmed $0.40 per million input and $1.50 per million output tokens, 50% off at launch. Inception
  • Claimed First answer token in under 320 milliseconds (median). Inception
Sources: Inception

Models & releases

Today in the September 30, 2026 edition · front page