Connect with us

NEWS

Google Keeps Speech Local as Microsoft Cuts the Rate

Google AI Edge Eloquent keeps Mac dictation on-device, while Microsoft’s MAI-Transcribe-2 sells cloud speech at $0.10 an hour and squeezes paid apps.

Published

on

Microsoft cut cloud speech-to-text to $0.10 per hour on September 3, three months after Google put a free dictation app on the Mac that never sends audio out. The two moves were sold as one local-AI wave. They are two different bets on the same microphone.

Google AI Edge Eloquent keeps the recording on the laptop. MAI-Transcribe-2 sells the hour so cheaply that Teams, Copilot, and contact-center audio have little reason to leave Microsoft’s stack. Paid dictation apps now sit in the gap.

A Free Mac App That Never Leaves the Laptop

On June 3, Google shipped Gemma 4 12B and a macOS build of Eloquent on the same day. The model is a 12-billion-parameter, encoder-free multimodal system with native audio input, released under Apache 2.0, and Google says Gemma 4 downloads have passed 150 million. It is built to 16GB of VRAM or unified memory, which puts it on ordinary Apple silicon laptops rather than a rack of GPUs.

The Mac app is the part most people will actually touch. Google’s AI Edge team calls Eloquent an on-device voice dictation app for Mac that runs the full feature set without a network link. A customizable hotkey dumps speech into whatever app is in front, and the same stack can transcribe audio or video files that never leave the disk.

Cleanup is the product, not a raw transcript. The app strips filler, offers writing styles, and accepts custom words so names and jargon stop getting “corrected.” Voice Edit, powered by Gemma 4 12B, lets you highlight a passage and speak an instruction at it. Google says that jump in instruction-following is a 60%+ quality lift over the prior on-device models.

Restructure these notes into an executive summary.

Google AI Edge team, Voice Edit example, June 3 developer post

Another sample command is “translate this into Hindi,” spoken at selected text rather than typed into a prompt box. That is dictation behaving like a local editor, which is why a free Google app on macOS is a problem for every subscription that charges for the same loop.

iPhone users got Eloquent first, in April, and that build can still send cleanup to Gemini unless cloud mode is off. The Mac version is the strict one. Google has not shipped a Windows twin, and a desktop Android build is still missing, so the offline pitch is strongest on Apple hardware.

Microsoft Priced the Microphone at 10 Cents

Microsoft AI, led by Mustafa Suleyman, answered on June 2 with seven in-house MAI models spanning image, voice, transcription, coding, and reasoning. The speech flagship then was MAI-Transcribe-1.5, which Suleyman called the best transcription model in the world, with state-of-the-art accuracy across 43 languages and up to five times the speed of Gemini 3.1, Scribe v2, and GPT-4o-Transcribe on long audio.

That model is not a laptop download. It is a Foundry API, with hooks into Copilot, Teams, GitHub, and Dynamics 365 Contact Center. Microsoft said it can transcribe an hour of audio in under 15 seconds and that keyword biasing cuts word error by up to 30% on FLEURS when a domain vocabulary is supplied. Independent bench site Artificial Analysis put it at 2.4% AA-WER, third overall, and very fast.

MAI-Voice-2 landed beside it as the pair’s mouth: 15 languages, emotion tags, and a 5 to 60 second clip to prompt a voice, preferred 72% of the time over MAI-Voice-1 in Microsoft’s tests. MAI-Thinking-1 filled the reasoning slot. The family was built to live inside Microsoft products, not to sit in a menu bar on a MacBook.

The two firms had already taken opposite paths on speech AI in June, and September pushed Microsoft further into the API. MAI-Transcribe-2, announced September 3, expands coverage to 60 languages, adds speaker diarization, word-level timestamps, verbatim or clean styles, code switching, and automatic language ID, and takes first place on FLEURS at a 5.2% average word error rate. Artificial Analysis now ranks it second at 2.0% AA-WER.

The number that will move buyers is the rate. Microsoft is offering the model at $0.10 per hour of audio as a limited-time price through the end of 2026, down 72% from the $0.36 an hour charged for MAI-Transcribe-1.5. Evals cited by Microsoft put it at 10 times the speed of OpenAI’s GPT-Transcribe, seven times ElevenLabs’ Scribe v2, and five times Google’s Gemini 3.5 Transcribe.

THE SPEECH STACK FROM SPRING TO SEPTEMBER

  1. Spring 2026: MAI-Transcribe-1 arrives on Foundry with 25 languages at $0.36 an hour.
  2. June 2, 2026: Microsoft AI launches seven models; MAI-Transcribe-1.5 adds 18 languages, to 43, and claims up to 5x speed on long audio.
  3. June 3, 2026: Google releases Gemma 4 12B and the fully local macOS Eloquent app.
  4. August 17, 2026: Wispr closes a $2 billion Series B on the back of paid dictation.
  5. September 3, 2026: MAI-Transcribe-2 ships at $0.10 an hour with 60 languages, diarization, and timestamps.

Spring’s $0.36 hour was already a hyperscaler rate. The September cut makes the same work cheaper than most specialist APIs without asking the customer to install a 12-billion-parameter model.

WHAT MAI-TRANSCRIBE-2 ADDS ON TOP OF 1.5

  • Speaker labels: Diarization attributes words to the right person inside one recording.
  • Word timestamps: Each word gets a time mark for captions, search, and editing.
  • Two styles: Verbatim keeps fillers for compliance; clean drops them for readable notes.
  • Code switching: Mixed pairs such as Hinglish and Spanglish stay in one pass.

Those are meeting-notes features, which is the point. Microsoft is not trying to win a Mac menu-bar war. It is trying to make the transcript that already happens inside its apps too cheap to rip out.

The $2 Billion Company in the Middle

Wispr, the San Francisco maker of Wispr Flow, raised $280 million in Series B funding on August 17 at a $2 billion valuation, led by Menlo Ventures, bringing total capital to $361 million. CEO and co-founder Tanay Kothari wrote that people have now produced more than 60 billion words with Flow, and that it is used at almost all Fortune 500 companies and more than 10,000 enterprises.

Seventeen days later, Microsoft’s hour cost $0.10. Google’s Mac app still costs nothing. That is a brutal sandwich for a company whose pitch is polished speech in every text field.

WISPR FLOW AFTER THE SERIES B

  • The round: $280 million at a $2 billion valuation, $361 million raised in total.
  • The usage: More than 60 billion words written with Flow, plus 10,000-plus companies.
  • The model: Canto, a first-party speech model in preview, aimed at noise, wind, accents, and music.
  • The edit cut: Wispr expects 30 to 35% fewer dictations that need a human fix in everyday use.

Kothari is not selling a quiet-room demo. He is selling the ugly cases, because that is where paid apps still beat a generic API.

We built this model for where people actually use Flow. In the hardest conditions, with background noise, wind, heavy accents or music, error rates fall from more than 30% of words to somewhere between 5 and 10%.

Tanay Kothari, CEO and co-founder, Wispr Flow blog, August 17, 2026

He also said the internal scoreboard is zero edit rate, the share of utterances that come back untouched. If Google’s free on-device cleanup and Microsoft’s 10-cent hour both get close enough, Wispr has to win on the last few points of accuracy, on hold-to-talk feel, and on meetings, where it already launched a note-taker. Mordor Intelligence, in a January outlook, put the speech recognition market at about $22.5 billion in 2026 and $61.7 billion by 2031, a 22.4% compound rate. There is room. There is also a free Google client and a Microsoft rate that looks like a loss leader.

600 Million Parameters and a 0.4-Second Loop

The same stretch of calendar filled in under the giants. NVIDIA released Nemotron 3.5 ASR on June 4 and June 11 as a 600-million-parameter streaming speech model that covers 40 language-locales from one checkpoint, with open weights and an OpenMDW-1.1 license. Chunk size is a dial: 80 milliseconds for agents, 1.12 seconds when accuracy matters more than lag. Teams that do not want audio on anyone else’s GPU can host it themselves, which is the actual local option for Windows shops Google never served.

AssemblyAI’s Universal-3 Pro, in a product update on mixed-language audio, posted a roughly 19% relative word-error drop on code-switching benches, with gains on CommonVoice and FLEURS sets and on Spanglish. That is still a cloud API. It is also the kind of specialist claim Microsoft’s 10-cent hour is designed to crowd.

Researchers at NTU, NUS, and CUHK released Audio-Interaction, a 3-billion-parameter streaming model that hears in 0.4-second slices and emits a silent-or-respond token after each slice. It scored 58.15 on MMAU under audio instructions, Apache-2.0 weights sit on Hugging Face, and the point of the design is continuous listening without a wake word. Always-on audio is the other end of Eloquent’s hold-a-hotkey model, and it is the version enterprises will fight over on privacy grounds.

WHERE THE NEW SPEECH STACK RUNS

Product Where it runs Scale Standout figure
Google AI Edge Eloquent / Gemma 4 12B On-device on Mac (iOS with optional cloud cleanup) 12B parameters, 16GB memory 100% on-device feature set on macOS
MAI-Transcribe-2 Microsoft Foundry, MAI Playground, OpenRouter 60 languages $0.10 per hour; 5.2% FLEURS WER
NVIDIA Nemotron 3.5 ASR Self-hosted open weights 600M parameters, 40 locales 80 ms to 1.12 s chunks
AssemblyAI Universal-3 Pro Cloud API Multilingual, code-switching ~19% relative WER drop on code-switch tests
Audio-Interaction Open weights, streaming loop 3B parameters 0.4 s decide step; MMAU 58.15

Self-hosting NVIDIA’s 600-million-parameter streamer is how a bank or a hospital gets local speech without waiting for Google to care about Windows. The 12-billion-parameter Gemma path is nicer to use and much hungrier for memory. Those are not the same product, even if both get filed under on-device.

Why Windows Users Still Buy Dictation Apps

Eloquent’s Mac-only desktop launch left the largest work PC market to everyone else. DictaFlow, a Windows-first hold-to-talk app, types into Remote Desktop and Citrix instead of pasting through a clipboard that locked-down desktops often block. It sells at $7 a month after a 2,000-word free tier, and the hold key is Ctrl+Windows by default. That is a narrow, ugly problem Google’s polished Mac client never had to solve.

Wispr Flow already runs on Windows as well as Mac and iPhone, which is part of what the $2 billion round is paying for. Windows 11 also ships Voice Access, which can dictate and control the PC without an account, and that baseline keeps getting better. None of those tools is Gemma 4 12B sitting in unified memory, rewriting a paragraph because you said so out loud.

So the Windows dictation market is still a pile of hotkeys, VDI workarounds, and hybrid local-plus-cloud engines. Microsoft could have put a first-party Eloquent rival in the Microsoft Store on June 2. It put MAI-Transcribe into Foundry instead, then cut the hour to a dime. Windows users who want the Google-style offline editor still buy a specialist or wait.

Foundry Turns Cheap Transcripts Into Capture

Top APIs are already bunched. Artificial Analysis has MAI-Transcribe-2 at 2.0% AA-WER and the previous MAI model at 2.4%, with several rivals inside a similar band. When error rates pack that tight, the remaining levers are speed, price, and which product already has the microphone.

Microsoft can sell a 10-cent hour because the transcript is how audio enters software it already runs. Teams records the meeting. Copilot sits on the same tenant. Dynamics 365 Contact Center is already wearing the headset. GitHub is already where the engineering talk lands. Cheap speech-to-text is the door, not the store. Early Transcribe-2 checks on September 9 found Chinese dialects coming through clearly even where the public language list was fuzzy about mainland Chinese varieties, which is the sort of field note that matters more than another leaderboard screenshot.

The same week, Meta put out Muse Voice Transcribe and the rest of the realtime stack kept moving, which is what a commodity looks like from the inside. Google’s counter is still the one that never bills an hour: a free Mac app whose audio never leaves the chassis, backed by a 12-billion-parameter model you can also load in LM Studio, Ollama, or Google AI Edge Gallery. That is a developer beachhead on Apple silicon, and it is a warning shot at every app that charged for “AI dictation” when the cleanup was the only scarce part.

MAI-Transcribe-2 is in Foundry, the MAI Playground, and OpenRouter at $0.10 an hour through the end of 2026. Eloquent remains a free Mac download that Google describes as fully offline across the whole feature set. They do not compete on the same machine, and they both squeeze the people who do.

Harry edits WinAddons, an independent news site that he owns and runs, covering Windows, Xbox, Azure, Microsoft 365, Teams, OneDrive, Outlook, the software built around them and Microsoft's business. His method comes from ten years in journalism, a reporter's years followed by an editor's, and the bulk of that decade has been spent watching Microsoft ship. His reporting starts with what Microsoft publishes: release notes and KB articles read in full, build numbers checked on an installed machine, MSRC advisories and the CVE records behind them, the Azure status history, lifecycle pages, store listings in the market they apply to, and the earnings releases and filings that carry the company's numbers. Every figure is checked against its source before publication, and a public corrections policy explains how mistakes are fixed and labelled. On security stories he does not publish exploit details before a fix is available, reporting what is affected and what to do instead. Pre-release features are labelled by channel and build, and a rumour is called a rumour. Readers can reach Harry at support@winaddons.com.

Continue Reading
Click to comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Trending