← AI Switchboard
AI Switchboardby Waggle
MODELS · September 24, 2026
Sep 23

The voice API price war arrives: Qwen cuts up to 95%, Google ships designed voices

Alibaba's Qwen shipped five audio models and cut speech prices by up to 95% on the same day Google launched Gemini 3.8 text-to-speech with voices built from written descriptions.

Two announcements landed on Wednesday that together reprice machine speech. Qwen released Audio-3.1 as five models: upgraded ASR, TTS and Realtime, plus two new ones, TTS-Next for audio creation and ASR-Next for audio understanding, described by Qwen as “one complete audio stack: understanding, generation, interaction & creation”. The price cuts are roughly 70% for text-to-speech, about 85% for realtime, and up to 95% for speech recognition.

Google, the same day, launched Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS. The headline capability is promptable voice design: rather than picking from a list, you describe the voice you want — role, accent, characteristics — in natural language, across more than 100 languages and dialects. There is also a library of over 2,000 production-ready voices and replication of a voice from a 30-second sample. Google reports itself top of Hume AI's Voice Design Benchmark at 71.4 and top on accent modelling at 60.8; those are Google's own readings of a third party's benchmark.

Every clip carries a watermark. “Every audio clip generated by our Gemini Audio models is watermarked with SynthID,” Google writes. “This imperceptible watermark is woven directly into the audio output.” Thirty-second voice cloning shipping with mandatory watermarking is the compromise position the industry has landed on, and it only works while the cheap alternatives do the same — which is exactly what an up-to-95% price cut from a competitor puts under pressure.

One figure widely repeated this week does not hold up: a $6.40 per million audio-input tokens price for Realtime-Plus. Alibaba Cloud's model catalogue lists the model without a price, and no primary source carries that number.

  • Confirmed Qwen Audio-3.1 ships five models with price cuts of roughly 70% (TTS), 85% (Realtime) and up to 95% (ASR). Qwen, via The Decoder
  • Confirmed Gemini 3.8 Flash TTS and Flash-Lite TTS add promptable voice design across 100+ languages, 2,000+ ready voices, and voice replication from a 30-second sample, all watermarked with SynthID. Google
  • Claimed Google places itself first on Hume AI's Voice Design Benchmark at 71.4 and first on accent modelling at 60.8 — its own reported results on a third party's benchmark. Google
Sources: Google · The Decoder

Models & releasesMoney & markets

This week in the September 24, 2026 edition · front page