Microsoft’s new voice transcription model beats Google and OpenAI on price and speed

Written by Jason Miller

Microsoft AI launched MAI-Transcribe-2the new voice recognition model that the Redmond company defines as the fastest, most accurate and economical on the market. The debut comes with an introductory price of just $0.10 per hour of audiovalid until the end of 2026.

The direct comparison is with the bulkier competitors in the industry: Google’s Gemini 3.5 Transcribe and OpenAI’s GPT-Transcribe. According to data released by Microsoft, processed together with the independent benchmark Artificial Analysis, MAI-Transcribe-2 processes audio up to 10 times faster of GPT-Transcribe, 7 times faster than Scribe v2 by ElevenLabs and 5 times faster than Gemini 3.5 Transcribe.

In terms of accuracy, the model ranks first in the benchmark FLEURSthe reference for multilingual transcription, covering 60 languages ​​with an average error rate of 5.2%. In the general ranking of Artificial Analysis, however, MAI-Transcribe-2 stops in second place for pure accuracy, despite defining what Microsoft calls the Pareto Frontier between precision and latency.

In addition to speed, MAI-Transcribe-2 introduces functions designed for professional use: the diarization of speakerswhich attributes each audio segment to a distinct voice, returning a structured dialogue instead of an undifferentiated block of text, and word-level timestamp. To complete the package there is keyword biasing for sector terminology, transcription modes configurable between verbatim and clean, and the management of code-switching between different languages ​​in the same conversation, a detail that comes in handy with Hindi-English or Spanish-English.

The system also integrates automatic language recognition and background noise management designed for real-world scenarios, from calls in open environments to recordings of meetings with multiple microphones. Microsoft points out that the reduced latency weighs especially on long audio, where competing models tend to slow down as the recording length increases.

The model, as announced by Microsoft AI, is already available in preview on Microsoft Foundry, on the MAI Playground platform and via Open Router.

The price cut comes at a time when the automatic transcription market has become crowded. Google, OpenAI and ElevenLabs already offer comparable services, but no one has so far pushed the cost per hour of audio so low, and the timed promotion suggests that the list price will return higher once 2026 ends.

In recent months, Microsoft AI has sequentially released models for transcription, speech synthesis and image generation: MAI-Transcribe-2 arrives after the first generation MAI-Transcribe-1, presented a few months earlier, confirming a pace of development that aims to make the company less dependent on OpenAI models for these specific functions.

Jason Miller

I'm Jason Miller, and I've been passionate about technology and storytelling for over a decade. As a lead writer at Herald Editorials, I strive to bring clarity and creativity to complex tech topics. When I'm not writing, you'll find me exploring the latest gadgets or hiking in the great outdoors.