Models

Microsoft MAI-Transcribe-2: Speech to Text at 10 Cents an Hour

Microsoft's new speech model costs 10 cents an hour, 72% below its predecessor. It tops FLEURS across 60 languages and runs at 410x real time. Per-language accuracy is not yet published.

Faruk TalmaçSeptember 5, 20263 min read2 views
Microsoft MAI-Transcribe-2: Speech to Text at 10 Cents an Hour

Ten cents for an hour of audio. That is the price Microsoft AI attached to MAI-Transcribe-2 when it shipped on September 3. The previous version, MAI-Transcribe-1.5, cost 36 cents an hour, so the cut is 72%. The rate is a launch promotion running through the end of 2026; what follows has not been announced.

Accuracy and speed, by the numbers

On Microsoft's own measurements the model ranks first on the FLEURS multilingual benchmark across 60 languages with an average word error rate (WER) of 5.2%. In the English-focused comparison table the order is: MAI-Transcribe-2 at 2.0%, ElevenLabs Scribe v2 at 2.2%, the previous 1.5 release at 2.4%, Google Gemini 3.5 Transcribe at 2.6%, OpenAI GPT-Transcribe at 3.3%. On the independent Artificial Analysis leaderboard it places second on WER and defines the efficient frontier on the accuracy-versus-latency curve.

The speed gap is wider than the accuracy gap. Median throughput is 410.7 times real time; Gemini 3.5 Transcribe manages 89.9x, Scribe v2 54.7x, GPT-Transcribe 40x. In practice an hour-long recording comes back in under ten seconds.

Features

  • Speaker diarization (who spoke when) and word-level timestamps
  • Keyword biasing: supply product names or people's names in advance to improve recognition
  • Two output styles, "verbatim" (with the ums and ahs) and "clean"
  • Automatic language identification and mid-sentence code-switching (Hinglish, Spanglish)
  • Robustness on noisy audio

The model is available on Microsoft Foundry, the MAI Playground and OpenRouter. Secondary reports put the core team at around ten people; Microsoft's announcement does not give a number.

Third launch in eight days

This is the third major speech-recognition release in just over a week: Google announced Gemini 3.5 Transcribe on August 28 and Meta followed with Muse Voice Transcribe this week. All three are chasing the same buyer: organizations with call-center, meeting-notes and captioning volume. Prices falling to cents per hour is the direct result.

What we don't know about language coverage

The announcement says 60 languages but does not publish the list, and per-language WER figures are not broken out. If your workload is in anything other than English, the first job is to measure it yourself on a 30-minute sample of your own recordings before committing.

The cost arithmetic

Our math: a support team taking 200 calls a day at six minutes each produces roughly 400 hours of audio a month. Transcribing it costs $40 at the new rate, $144 at the old one, and 400 human hours if someone listens. Once transcription approaches zero, the question stops being "should we transcribe" and becomes "which model analyzes the transcripts," and that is where the real bill forms. The 2027 price is unknown, so we would not build a long-term cost model on the promotional rate. And for any business recording customers: check that your consent notices cover sending voice data to a third-party cloud before you switch this on.

Sources: Microsoft AI announcement, eWeek: comparison table, Martin Cid Magazine

Share This Article

Faruk Talmaç

Written by

Faruk Talmaç

Co-Founder & Editor

Co-founder of YZ Uzman, with 20+ years of experience in web design and software development.

More news

Want to put this technology to work in your business?

Let's talk