Connect with us

NEWS

Microsoft AI Was Set Free to Build Its Own Frontier

Suleyman said Microsoft AI was set free from OpenAI to build superintelligence with MAI models, Maia chips, and customer work data.

Published

on

Mustafa Suleyman said at Microsoft Build on June 2, 2026 that his lab had been set free from OpenAI to pursue superintelligence. The CEO of Microsoft AI timed the contract change at about six months earlier. The same day, the company launched seven in-house MAI models.

Those models are a first public proof, not the finished lab. The lasting change is a loop of customer work traces, first-party models, and Maia chips that lets Microsoft treat OpenAI as one option inside Foundry rather than the only path.

Microsoft Could Not Train Past a FLOPS Cap

For years the OpenAI partnership, which began in 2019 and grew past a $13 billion investment, gave Microsoft early access to frontier models while fencing off its own AGI work. OpenAI would train the biggest systems. Microsoft would host them on Azure and fold them into Copilot. The deal even capped how large a model Microsoft itself could train, using a compute ceiling measured in FLOPS.

That fence came down in writing on October 28, 2025. OpenAI’s own partnership note said Microsoft can now independently pursue AGI alone or with third parties. Microsoft’s stake in OpenAI Group PBC was valued at about $135 billion, or roughly 27 percent. Intellectual-property rights for models and products run through 2032, including post-AGI systems with safety limits. OpenAI also contracted to buy an extra $250 billion of Azure services, while Microsoft lost its right of first refusal to supply all of OpenAI’s compute.

One limit remains if Microsoft uses OpenAI’s IP to chase AGI before an independent panel verifies it: compute thresholds still apply, though OpenAI said those thresholds sit well above the systems used to train leading models now. MAI-Thinking-1 sidesteps that clause because Suleyman’s team says it trained the model from scratch on licensed data, without distilling another lab’s weights.

FROM THE CONTRACT TO THE LAB

  1. October 28, 2025: OpenAI and Microsoft sign a new agreement that lets Microsoft pursue AGI on its own.
  2. November 6, 2025: Suleyman forms the MAI Superintelligence Team and names the goal humanist superintelligence.
  3. January 26, 2026: Microsoft puts Maia 200, its second-generation AI accelerator, into production in Iowa.
  4. June 2, 2026: At Build, the lab ships seven MAI models and Suleyman says the contract set the team free about six months earlier.
  5. July 29, 2026: Suleyman reports MAI models already cutting GPU cost inside Copilot, PowerPoint, OneDrive, and Dynamics 365.

The October paper still calls OpenAI Microsoft’s frontier model partner. Suleyman did not describe a breakup. He described room to build.

Seven Models, One Lab, Six Months

Suleyman’s Build post introduced a family of seven new models spanning reasoning, code, images, transcription, and voice. He wrote that the lab does not distill from other labs and does not rely on opaque data. The pre-training mix for the reasoning line, he said in the Build interview, was about 50 percent high-quality code, with the rest commercially licensed and curated.

Microsoft AI posted the flagship’s scorecard the next day. MAI-Thinking-1 is a mixture-of-experts model with 35 billion active parameters and 1 trillion total. The company said it scored 52.8 percent on SWE-Bench Pro, hit 97 percent on AIME 2025, and was preferred to Sonnet 4.6 in blind side-by-side tests. Suleyman also pointed developers to a 109-page technical paper for the details that did not fit the keynote.

THE MAI FAMILY AT BUILD

Model Job Figure Microsoft cited
MAI-Thinking-1 Reasoning and software engineering 35 billion active parameters, 52.8 percent on SWE-Bench Pro
MAI-Code-1-Flash Agentic coding in GitHub Copilot and VS Code 5 billion active parameters
MAI-Image-2.5 Text-to-image and image editing No. 2 on the Arena image boards
MAI-Transcribe-1.5 Speech to text 43 languages, 5 times faster than rival models
MAI-Voice-2 Speech generation 15 languages, with a Flash variant still coming at launch

Flash variants complete the seven-model count. For the first time, Microsoft said developers can tune MAI weights themselves on OpenRouter, Fireworks, and Baseten, besides hosting them on Foundry. Suleyman still called the drop a start. “Our job is to make sure that when we look out to 2030 and beyond, we have the capacity not just to buy models from third parties, but to build the absolute frontier, the best models in the world,” he said. “That’s a long transition.”

Inside Mayo Clinic and the Excel Gyms

The product that turns those models into a moat is Frontier Tuning. Microsoft runs reinforcement-learning environments it calls training gyms, so an agent can practice a company’s real tasks without touching production systems. The valuable data, Suleyman said, is no longer the open web. “We’ve sort of hoovered up all of the obvious pools of training data,” he said. The next pool is the trace of work inside Excel, Word, Teams, Jira, and CRM tools.

He put a number on the installed base: Microsoft serves 493 of the Fortune 500 through Azure. Frontier Tuning is how that seat becomes model quality. A MAI model tuned for Excel, the lab said, matches GPT 5.4 while running up to 10 times more efficient. In the Build keynote, Suleyman said a McKinsey tuning run beat GPT-5.5 on quality at 10 times lower cost.

EARLY FRONTIER TUNING PARTNERS

  • Mayo Clinic: Co-creates a healthcare frontier model on de-identified clinical data; Mayo owns the model and runs it first in its own environment before Foundry.
  • EY: Tunes a tax-advisory agent for 75,000 professionals.
  • Land O’Lakes: Reported better grounded outputs and style compliance after tuning.
  • Pearson: Uses tuned models for learning-science feedback in Communication Coach.

Mayo’s ownership clause is the tell. Microsoft is not scraping a hospital’s notes into a shared brain. It is selling the gym, the silicon, and the base model, then leaving the tuned weights with the customer. That is a very different business from renting GPT through Azure and hoping the contract holds.

Maia 200 Underprices a GB200 by 30 Percent

None of the gyms work if tokens stay this expensive. Suleyman said Microsoft is the largest buyer of GPUs on the planet, and the largest buyer of Nvidia GB200s and GB300s, which is why Microsoft’s rising AI hardware bill now shows up in supplier margins as clearly as it does in Azure capex. He also said the company will keep buying Nvidia “for many, many years to come.”

In parallel it is shipping its own chip. On January 26, 2026, Scott Guthrie, executive vice president for Cloud and AI, said Maia 200 is built on TSMC’s 3nm process with 216GB of HBM3e at 7 TB/s and more than 140 billion transistors. He put each chip at over 10 petaFLOPS in 4-bit precision inside a 750W envelope. The first racks sit in Azure’s US Central region near Des Moines, Iowa, with US West 3 near Phoenix, Arizona next. The same fleet still serves OpenAI’s GPT-5.2 models. The Superintelligence team also uses Maia 200 to generate synthetic data for its own training runs.

Suleyman’s cost claim is sharper than the spec sheet. He said Maia 200 is 30 percent cheaper than a GB200 inside Microsoft’s clusters. When the lab co-designs MAI models for that chip, he later wrote, it sees 40 percent better performance per watt. That is the closed loop in hardware form: own model, own accelerator, own cloud, tuned on the customer’s tasks.

What Set Free From OpenAI Changed

Suleyman delivered the line backstage at Fort Mason during Build, without raising his voice.

We were only sort of set free from our contract with OpenAI about six months ago to formally pursue superintelligence. So this is very early days.

Mustafa Suleyman, CEO of Microsoft AI, at Microsoft Build 2026

He also said there is no urgent hole to fill in three months or six. “We have OpenAI, we have Anthropic, we have thousands of models inside Foundry. So there’s already a huge amount of optionality available to us.” The October contract did not kick OpenAI out. It stopped Microsoft from being barred from growing a second brain.

That mix is the point of Microsoft building its own AI while still listing Anthropic and OpenAI in the same catalog. Foundry stays a mall. MAI is the store Microsoft can restock if a partner raises prices, hits a policy wall, or simply goes elsewhere for compute. Suleyman put the mission in five-year language: produce state-of-the-art frontier-scale models, not because Copilot dies tomorrow without them, but because a platform company cannot rent its core layer forever.

He also rejected the idea that models are turning into identical cheap tokens. Distinct lineages will reflect different training goals, he said, and quality tokens will matter more than brute-force scale. A coding-and-agents lineage trained on licensed data is, in that telling, a different product from a consumer chat model trained on the open web. Whether buyers pay extra for that lineage is the bet.

Superintelligence, Scoped to Human Problems

On November 6, 2025, the day the Superintelligence Team went public, Suleyman framed the work as humanist superintelligence: advanced systems that stay tools, shaped by human intent and subordinate to human goals. He drew a hard line against an unbounded, highly autonomous entity. The lab, he said, should be problem-oriented and domain-specific, with real limits on autonomy.

https://x.com/mustafasuleyman/status/1986433769046483430

That post invited a familiar jab, including a reply asking whether Microsoft would become a nonprofit to match the language. It will not. Mayo’s model, EY’s tax agent, and GitHub Copilot are commercial Azure products. The “humanist” label is a design constraint Suleyman wants on the record: medical and workflow systems that stay controllable, not a lab racing to outgrow its operators. He has said some other labs treat a smarter-than-all-of-us system as both inevitable and desirable. His rule is the opposite. Build it only if it stays subordinate.

He knows how long a lab actually takes. He co-founded DeepMind in 2010, and that group spent years caught between research and product. “If you rush it, you’ll screw it up,” he said at Build.

Copilot and PowerPoint Already Run MAI Models

The rush test arrived sooner than the 2030 slogan. On July 29, 2026, Suleyman wrote that the hill-climbing machine had spent the quarter on token efficiency, not bigger generalist scores, and that Microsoft had shipped more than a dozen new models across image, voice, transcription, coding, and security since the prior quarter.

WHERE MAI ALREADY CUTS GPU COST

  • GitHub Copilot: Millions of developers have used MAI-Code-1-Flash since June, with a 10 percent higher code accept rate and 10 percent lower median token use than GPT-5.4 Mini and Claude Haiku 4.5 in VS Code.
  • PowerPoint and OneDrive: MAI-Image-2.5-Flash is the default in key flows, cutting GPU cost up to 84 percent versus GPT-Image-2 and lifting OneDrive save rates 26 percent.
  • Dynamics 365: MAI-Voice-2-Flash now powers Contact Center agents for customers such as T-Mobile and EasyJet, cutting GPU cost up to 89 percent.
  • Dragon Copilot: MAI-Transcribe-1.5 runs a 58-language workflow used by 170,000 medical providers and 28 million encounters last quarter, with a 50 percent relative drop in transcription and language-ID errors in Microsoft’s tests.

A security-tuned sibling, MAI-Cyber-1-Flash, hit 96 percent on CyberGym at 50 percent of the cost of Microsoft’s prior MDASH offering, on H100s. That system is built to handle up to 90 percent of tasks so the fleet can save GPT 5.4 for the 10 percent of problems that still need it. On September 4, 2026, MAI-Image-2.6 took No. 2 on Arena for both text-to-image and editing, with a Flash variant Microsoft said is 2.8 times faster than GPT-Image-2-Medium and 72 percent more efficient.

The pattern is not a clean swap of OpenAI for MAI. It is a cheaper default for the common case, with a partner model still on the hook for the hard slice. Suleyman wrote that every model in a product should be substitutable, because any one lab can vanish behind a security incident, a policy fight, or a geopolitical split. That sentence is the whole strategy, minus the keynote lighting.

The sticker on his laptop at Build read “Patience and urgency.” The July cost cuts are the urgency. The five-year clock to a self-sufficient frontier lab is the patience, and it has already started billing in Copilot tokens rather than in slogans.

Harry edits WinAddons, an independent news site that he owns and runs, covering Windows, Xbox, Azure, Microsoft 365, Teams, OneDrive, Outlook, the software built around them and Microsoft's business. His method comes from ten years in journalism, a reporter's years followed by an editor's, and the bulk of that decade has been spent watching Microsoft ship. His reporting starts with what Microsoft publishes: release notes and KB articles read in full, build numbers checked on an installed machine, MSRC advisories and the CVE records behind them, the Azure status history, lifecycle pages, store listings in the market they apply to, and the earnings releases and filings that carry the company's numbers. Every figure is checked against its source before publication, and a public corrections policy explains how mistakes are fixed and labelled. On security stories he does not publish exploit details before a fix is available, reporting what is affected and what to do instead. Pre-release features are labelled by channel and build, and a rumour is called a rumour. Readers can reach Harry at support@winaddons.com.

Continue Reading
Click to comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Trending