NEWS
Google Keeps Speech Local as Microsoft Cuts the Rate
Google AI Edge Eloquent keeps Mac dictation on-device, while Microsoft’s MAI-Transcribe-2 sells cloud speech at $0.10 an hour and squeezes paid apps.
Microsoft cut cloud speech-to-text to $0.10 per hour on September 3, three months after Google put a free dictation app on the Mac that never sends audio out. The two moves were sold as one local-AI wave. They are two different bets on the same microphone.
Google AI Edge Eloquent keeps the recording on the laptop. MAI-Transcribe-2 sells the hour so cheaply that Teams, Copilot, and contact-center audio have little reason to leave Microsoft’s stack. Paid dictation apps now sit in the gap.
A Free Mac App That Never Leaves the Laptop
On June 3, Google shipped Gemma 4 12B and a macOS build of Eloquent on the same day. The model is a 12-billion-parameter, encoder-free multimodal system with native audio input, released under Apache 2.0, and Google says Gemma 4 downloads have passed 150 million. It is built to 16GB of VRAM or unified memory, which puts it on ordinary Apple silicon laptops rather than a rack of GPUs.
The Mac app is the part most people will actually touch. Google’s AI Edge team calls Eloquent an on-device voice dictation app for Mac that runs the full feature set without a network link. A customizable hotkey dumps speech into whatever app is in front, and the same stack can transcribe audio or video files that never leave the disk.
Cleanup is the product, not a raw transcript. The app strips filler, offers writing styles, and accepts custom words so names and jargon stop getting “corrected.” Voice Edit, powered by Gemma 4 12B, lets you highlight a passage and speak an instruction at it. Google says that jump in instruction-following is a 60%+ quality lift over the prior on-device models.
Restructure these notes into an executive summary.
Google AI Edge team, Voice Edit example, June 3 developer post
Another sample command is “translate this into Hindi,” spoken at selected text rather than typed into a prompt box. That is dictation behaving like a local editor, which is why a free Google app on macOS is a problem for every subscription that charges for the same loop.
iPhone users got Eloquent first, in April, and that build can still send cleanup to Gemini unless cloud mode is off. The Mac version is the strict one. Google has not shipped a Windows twin, and a desktop Android build is still missing, so the offline pitch is strongest on Apple hardware.
Microsoft Priced the Microphone at 10 Cents
Microsoft AI, led by Mustafa Suleyman, answered on June 2 with seven in-house MAI models spanning image, voice, transcription, coding, and reasoning. The speech flagship then was MAI-Transcribe-1.5, which Suleyman called the best transcription model in the world, with state-of-the-art accuracy across 43 languages and up to five times the speed of Gemini 3.1, Scribe v2, and GPT-4o-Transcribe on long audio.
That model is not a laptop download. It is a Foundry API, with hooks into Copilot, Teams, GitHub, and Dynamics 365 Contact Center. Microsoft said it can transcribe an hour of audio in under 15 seconds and that keyword biasing cuts word error by up to 30% on FLEURS when a domain vocabulary is supplied. Independent bench site Artificial Analysis put it at 2.4% AA-WER, third overall, and very fast.
MAI-Voice-2 landed beside it as the pair’s mouth: 15 languages, emotion tags, and a 5 to 60 second clip to prompt a voice, preferred 72% of the time over MAI-Voice-1 in Microsoft’s tests. MAI-Thinking-1 filled the reasoning slot. The family was built to live inside Microsoft products, not to sit in a menu bar on a MacBook.
The two firms had already taken opposite paths on speech AI in June, and September pushed Microsoft further into the API. MAI-Transcribe-2, announced September 3, expands coverage to 60 languages, adds speaker diarization, word-level timestamps, verbatim or clean styles, code switching, and automatic language ID, and takes first place on FLEURS at a 5.2% average word error rate. Artificial Analysis now ranks it second at 2.0% AA-WER.
The number that will move buyers is the rate. Microsoft is offering the model at $0.10 per hour of audio as a limited-time price through the end of 2026, down 72% from the $0.36 an hour charged for MAI-Transcribe-1.5. Evals cited by Microsoft put it at 10 times the speed of OpenAI’s GPT-Transcribe, seven times ElevenLabs’ Scribe v2, and five times Google’s Gemini 3.5 Transcribe.
THE SPEECH STACK FROM SPRING TO SEPTEMBER
- Spring 2026: MAI-Transcribe-1 arrives on Foundry with 25 languages at $0.36 an hour.
- June 2, 2026: Microsoft AI launches seven models; MAI-Transcribe-1.5 adds 18 languages, to 43, and claims up to 5x speed on long audio.
- June 3, 2026: Google releases Gemma 4 12B and the fully local macOS Eloquent app.
- August 17, 2026: Wispr closes a $2 billion Series B on the back of paid dictation.
- September 3, 2026: MAI-Transcribe-2 ships at $0.10 an hour with 60 languages, diarization, and timestamps.
Spring’s $0.36 hour was already a hyperscaler rate. The September cut makes the same work cheaper than most specialist APIs without asking the customer to install a 12-billion-parameter model.
WHAT MAI-TRANSCRIBE-2 ADDS ON TOP OF 1.5
- Speaker labels: Diarization attributes words to the right person inside one recording.
- Word timestamps: Each word gets a time mark for captions, search, and editing.
- Two styles: Verbatim keeps fillers for compliance; clean drops them for readable notes.
- Code switching: Mixed pairs such as Hinglish and Spanglish stay in one pass.
Those are meeting-notes features, which is the point. Microsoft is not trying to win a Mac menu-bar war. It is trying to make the transcript that already happens inside its apps too cheap to rip out.
The $2 Billion Company in the Middle
Wispr, the San Francisco maker of Wispr Flow, raised $280 million in Series B funding on August 17 at a $2 billion valuation, led by Menlo Ventures, bringing total capital to $361 million. CEO and co-founder Tanay Kothari wrote that people have now produced more than 60 billion words with Flow, and that it is used at almost all Fortune 500 companies and more than 10,000 enterprises.
Seventeen days later, Microsoft’s hour cost $0.10. Google’s Mac app still costs nothing. That is a brutal sandwich for a company whose pitch is polished speech in every text field.
WISPR FLOW AFTER THE SERIES B
- The round: $280 million at a $2 billion valuation, $361 million raised in total.
- The usage: More than 60 billion words written with Flow, plus 10,000-plus companies.
- The model: Canto, a first-party speech model in preview, aimed at noise, wind, accents, and music.
- The edit cut: Wispr expects 30 to 35% fewer dictations that need a human fix in everyday use.
Kothari is not selling a quiet-room demo. He is selling the ugly cases, because that is where paid apps still beat a generic API.
We built this model for where people actually use Flow. In the hardest conditions, with background noise, wind, heavy accents or music, error rates fall from more than 30% of words to somewhere between 5 and 10%.
Tanay Kothari, CEO and co-founder, Wispr Flow blog, August 17, 2026
He also said the internal scoreboard is zero edit rate, the share of utterances that come back untouched. If Google’s free on-device cleanup and Microsoft’s 10-cent hour both get close enough, Wispr has to win on the last few points of accuracy, on hold-to-talk feel, and on meetings, where it already launched a note-taker. Mordor Intelligence, in a January outlook, put the speech recognition market at about $22.5 billion in 2026 and $61.7 billion by 2031, a 22.4% compound rate. There is room. There is also a free Google client and a Microsoft rate that looks like a loss leader.
600 Million Parameters and a 0.4-Second Loop
The same stretch of calendar filled in under the giants. NVIDIA released Nemotron 3.5 ASR on June 4 and June 11 as a 600-million-parameter streaming speech model that covers 40 language-locales from one checkpoint, with open weights and an OpenMDW-1.1 license. Chunk size is a dial: 80 milliseconds for agents, 1.12 seconds when accuracy matters more than lag. Teams that do not want audio on anyone else’s GPU can host it themselves, which is the actual local option for Windows shops Google never served.
AssemblyAI’s Universal-3 Pro, in a product update on mixed-language audio, posted a roughly 19% relative word-error drop on code-switching benches, with gains on CommonVoice and FLEURS sets and on Spanglish. That is still a cloud API. It is also the kind of specialist claim Microsoft’s 10-cent hour is designed to crowd.
Researchers at NTU, NUS, and CUHK released Audio-Interaction, a 3-billion-parameter streaming model that hears in 0.4-second slices and emits a silent-or-respond token after each slice. It scored 58.15 on MMAU under audio instructions, Apache-2.0 weights sit on Hugging Face, and the point of the design is continuous listening without a wake word. Always-on audio is the other end of Eloquent’s hold-a-hotkey model, and it is the version enterprises will fight over on privacy grounds.
WHERE THE NEW SPEECH STACK RUNS
| Product | Where it runs | Scale | Standout figure |
|---|---|---|---|
| Google AI Edge Eloquent / Gemma 4 12B | On-device on Mac (iOS with optional cloud cleanup) | 12B parameters, 16GB memory | 100% on-device feature set on macOS |
| MAI-Transcribe-2 | Microsoft Foundry, MAI Playground, OpenRouter | 60 languages | $0.10 per hour; 5.2% FLEURS WER |
| NVIDIA Nemotron 3.5 ASR | Self-hosted open weights | 600M parameters, 40 locales | 80 ms to 1.12 s chunks |
| AssemblyAI Universal-3 Pro | Cloud API | Multilingual, code-switching | ~19% relative WER drop on code-switch tests |
| Audio-Interaction | Open weights, streaming loop | 3B parameters | 0.4 s decide step; MMAU 58.15 |
Self-hosting NVIDIA’s 600-million-parameter streamer is how a bank or a hospital gets local speech without waiting for Google to care about Windows. The 12-billion-parameter Gemma path is nicer to use and much hungrier for memory. Those are not the same product, even if both get filed under on-device.
Why Windows Users Still Buy Dictation Apps
Eloquent’s Mac-only desktop launch left the largest work PC market to everyone else. DictaFlow, a Windows-first hold-to-talk app, types into Remote Desktop and Citrix instead of pasting through a clipboard that locked-down desktops often block. It sells at $7 a month after a 2,000-word free tier, and the hold key is Ctrl+Windows by default. That is a narrow, ugly problem Google’s polished Mac client never had to solve.
Wispr Flow already runs on Windows as well as Mac and iPhone, which is part of what the $2 billion round is paying for. Windows 11 also ships Voice Access, which can dictate and control the PC without an account, and that baseline keeps getting better. None of those tools is Gemma 4 12B sitting in unified memory, rewriting a paragraph because you said so out loud.
So the Windows dictation market is still a pile of hotkeys, VDI workarounds, and hybrid local-plus-cloud engines. Microsoft could have put a first-party Eloquent rival in the Microsoft Store on June 2. It put MAI-Transcribe into Foundry instead, then cut the hour to a dime. Windows users who want the Google-style offline editor still buy a specialist or wait.
Foundry Turns Cheap Transcripts Into Capture
Top APIs are already bunched. Artificial Analysis has MAI-Transcribe-2 at 2.0% AA-WER and the previous MAI model at 2.4%, with several rivals inside a similar band. When error rates pack that tight, the remaining levers are speed, price, and which product already has the microphone.
Microsoft can sell a 10-cent hour because the transcript is how audio enters software it already runs. Teams records the meeting. Copilot sits on the same tenant. Dynamics 365 Contact Center is already wearing the headset. GitHub is already where the engineering talk lands. Cheap speech-to-text is the door, not the store. Early Transcribe-2 checks on September 9 found Chinese dialects coming through clearly even where the public language list was fuzzy about mainland Chinese varieties, which is the sort of field note that matters more than another leaderboard screenshot.
The same week, Meta put out Muse Voice Transcribe and the rest of the realtime stack kept moving, which is what a commodity looks like from the inside. Google’s counter is still the one that never bills an hour: a free Mac app whose audio never leaves the chassis, backed by a 12-billion-parameter model you can also load in LM Studio, Ollama, or Google AI Edge Gallery. That is a developer beachhead on Apple silicon, and it is a warning shot at every app that charged for “AI dictation” when the cleanup was the only scarce part.
MAI-Transcribe-2 is in Foundry, the MAI Playground, and OpenRouter at $0.10 an hour through the end of 2026. Eloquent remains a free Mac download that Google describes as fully offline across the whole feature set. They do not compete on the same machine, and they both squeeze the people who do.
-
NEWS4 months agoWarzone Leaves Xbox One and PS4 After Season 06
-
NEWS3 months agoMicrosoft AI Was Set Free to Build Its Own Frontier
-
MICROSOFT 3653 months agoMicrosoft IQ Turns Workplace Data Into a Metered Agent Brain
-
NEWS3 months agoXbox Games Showcase 2026 Split the Catalog in Two
-
MICROSOFT 3653 months agoNadella Banned Addiction Talk While Scout Kept Heartbeat
-
NEWS4 months agoModern Warfare 4 Splits Its Audience Before the October Launch
-
NEWS3 months agoInfinity Ward Bets Modern Warfare 4 on a Paid DMZ
-
NEWS3 months agoDragonwilds Hits Xbox, but Steam Saves Stay Put
