Connect with us

AZURE

AWS’s CloudFront Outage Repeats a Script Microsoft’s Azure Already Ran

AWS CloudFront’s three-hour outage on July 16 matches the same control-plane failure Microsoft’s Azure Front Door suffered twice in October 2025.

Published

on

AWS CloudFront broke for three hours and 33 minutes on July 16, and took Canvas, Blackboard and Hugging Face down with it. Amazon traced the failure to a single capacity limit inside one Frankfurt data center. Microsoft’s own cloud ran the identical failure twice in three weeks last October.

Azure Front Door failed the same way on October 9 and October 29, 2025, pulling down Microsoft 365, Outlook, Xbox Live and Copilot along with it. Microsoft published a detailed engineering post-mortem within weeks and froze its own configuration pipeline to fix it. Four days after its own outage, AWS still has not said what broke.

A Frankfurt Capacity Limit Cut Off Canvas, Blackboard and Hugging Face

At 3:45 a.m. ET on Thursday, July 16, AWS CloudFront began returning 5xx errors to every customer using its VPC Origins feature. The cascade ran for three hours and 33 minutes, hitting Canvas and Blackboard at universities across North America, machine learning platform Hugging Face, the UK National Lottery’s website, and networking hardware maker Ubiquiti’s cloud services.

AWS traced the root cause to an internal capacity limit inside a single availability zone, euc1-az2, part of its EU-CENTRAL-1 region in Frankfurt, Germany. When that limit was hit, the system responsible for pushing routing configuration to CloudFront’s edge network processors stopped loading updates. The edge servers kept running. They just had no valid instructions for routing traffic to VPC Origins backends, so every request came back as an error, regardless of where the customer or their users were located.

Other confirmed victims included collaboration tool Coda, telehealth provider Doxy, identity platform Frontegg, data infrastructure provider TigerData, Tailscale’s admin console, and Japanese platform Hatena Blog. Hugging Face, the machine learning developer platform that separately spent this year confronting malware hiding stolen data inside hosted models, was among the highest profile names caught up in the outage. AWS engineers identified the issue around 7:16 a.m. ET, and the resolution summary was posted at 8:21 a.m. ET.

One Feature, No Fallback

VPC Origins, announced at AWS re:Invent in November 2024, closed a real security gap. Before it existed, any load balancer or server sitting behind CloudFront needed a public IP address, and that IP was an attack surface. Anyone who found it could bypass CloudFront’s firewall rules and rate limiting entirely.

VPC Origins fixed that by routing CloudFront’s edge traffic through a service-managed network interface inside the customer’s own private subnet. No public IP. No internet gateway. Just CloudFront in, and nothing else.

That same design is what made Thursday’s failure total. With no public origin left standing, CloudFront was the only door in. When its connection layer broke, there was no side entrance to fall back to. AWS’s own suggested workaround, temporarily switching back to a public origin, only works if a team has already written and tested that change in advance. Most had not, at 3:45 in the morning.

Azure Front Door Ran the Same Script Twice Last October

Long before AWS customers were staring at 5xx errors, Microsoft’s equivalent service had already lived through this exact failure mode, not once but twice inside the same month.

The October 9 Warning Shot

Microsoft’s own account, published on its Azure Networking blog, describes a manual cleanup of stuck tenant metadata that bypassed Front Door’s configuration protection layer on October 9, 2025. The incompatible metadata reached a subset of edge sites and triggered a latent defect, causing availability impact Microsoft measured at roughly 6% of its Europe capacity and 16% of its Africa capacity. It was contained. It was also, in hindsight, a preview.

October 29 Goes Global

Three weeks later, a different sequence of configuration changes across two control-plane versions produced another batch of incompatible metadata. This time the protection system let it through before a delayed processing task crashed edge sites worldwide. The outage started around 15:41 UTC, roughly 11:40 a.m. ET, and knocked out Microsoft 365, Outlook, Xbox Live, Copilot, the Azure Portal, and identity service Entra ID. It landed hours before its own quarterly earnings release.

We are working to address an issue affecting Azure Front Door that is impacting the availability of some services.

A Microsoft spokesperson gave that statement to CNBC as the outage spread. Microsoft’s minute-by-minute timeline of the Front Door recovery shows engineers blocking all new configuration propagation within two minutes of the data-plane impact starting, then spending hours reloading configuration across the global fleet and rebalancing traffic to avoid overloading nodes as they came back.

How Do AWS, Azure and Google’s Recent Outages Compare?

Four major cloud incidents in just over a year share the same shape: a change or limit hit in one place, and a control plane that could not stop it from reaching everywhere. Root causes differ in the details but not in the structure, and only some of the companies involved have said exactly what happened.

Outage Date Root Cause Duration or Scope Key Systems Hit
AWS CloudFront (VPC Origins) July 16, 2026 Capacity limit in Frankfurt blocked routing updates 3 hours 33 minutes Canvas, Blackboard, Hugging Face, UK National Lottery
Azure Front Door (first incident) October 9, 2025 Manual override bypassed a config safeguard Regional: ~6% of Europe, ~16% of Africa capacity Contained to affected regional customers
Azure Front Door (second incident) October 29, 2025 Incompatible metadata across two control-plane versions About 7 hours to full recovery Microsoft 365, Outlook, Xbox Live, Copilot, Entra ID
Google Cloud Platform June 2025 Bad configuration update to authorization infrastructure Not publicly detailed Spotify and Cloudflare authentication

Cisco’s network intelligence unit tracked two configuration incidents three weeks apart and noted the October 9 event stayed geographically contained while the October 29 one did not, because the failure that time reached the global protection layer itself rather than a regional subset of edge sites.

Microsoft’s Post-Mortem Highlights AWS’s Silence

Here is where the two companies’ paths split. Microsoft’s Azure Networking team published a public, technical breakdown of both October incidents, named the specific defects behind each one, and said it deliberately blocked all Front Door configuration changes from October 29 to November 5 while it rushed out hardening fixes. The post is the first in a promised multi-part series on configuration resiliency, tenant isolation and recovery.

AWS has offered no equivalent. It confirmed the Frankfurt capacity limit and the euc1-az2 availability zone. It has not said what resource was actually running out.

  • What AWS has confirmed. The euc1-az2 availability zone hit an internal capacity constraint that stopped routing configuration from loading onto CloudFront’s edge network processors, producing global 5xx errors for VPC Origins customers for 3 hours 33 minutes.
  • What AWS has confirmed. All affected services were restored, and customers who switched to the public-origin workaround can safely revert to VPC Origins.
  • What AWS has not said. Whether the internal constraint tracked a connection count, a thread pool, a memory ceiling, or some other resource limit.
  • What AWS has not said. Whether that limit has been raised, or whether any circuit-breaker mechanism has been added to stop a repeat.

Mayur Upadhyaya, chief executive of API observability firm APIContext, told Cybernews the industry is “increasingly consolidating around a small number of providers because they’re the most convenient and economically attractive choice,” and that a fault that once hit a handful of organizations can now reach thousands at once because so many businesses lean on the same infrastructure. An estimated 94% of enterprise services worldwide depend on at least one of AWS, Azure or Google Cloud, and the three together hold more than 62% of the global cloud market, according to Rest of World’s analysis of 2025 outages.

What Azure and Microsoft 365 Admins Do Before the Next Outage

IncidentHub founder Hrishikesh Barua, whose monitoring service detected the July 16 cascade in real time, laid out preparation steps that apply just as much to Azure customers as to AWS ones. It is not exotic. All of it takes work done before an incident, not during one.

  • Pre-stage the failover. Whether the plan is switching a CloudFront origin to public or failing an application off Front Door, write and test the Terraform or equivalent change now, so it takes minutes to deploy instead of hours to improvise at 4 a.m.
  • Map second-order dependencies. Entra ID, Azure AD B2C and other identity services sit behind the same kind of global edge layer. An identity provider outage locks users out of applications that are running perfectly normally on their own.
  • Audit vendor CDN redundancy. A university running Canvas has no multi-CDN option for Canvas itself. Ask every SaaS vendor in the stack whether they have one, and weigh the answer into vendor risk reviews.

Region failover, the standard response to a zonal problem, does not help when the failure sits in a control plane that spans every region at once. That is what happened to Front Door on October 29, and it is what happened to CloudFront on July 16. Microsoft locked its own engineers out of Front Door’s configuration for a week to make sure it would not happen a third time. AWS, four days on, has not said what it is doing about Frankfurt.

Frequently Asked Questions

Did the AWS CloudFront outage affect Microsoft 365 or Azure directly?

No direct link has been reported. The July 16 outage was confined to AWS’s own CloudFront VPC Origins feature and its customers, such as Hugging Face and Ubiquiti. Microsoft 365 and Azure run on Microsoft’s own infrastructure and were not named among the affected services in any AWS or third-party incident tracking of this event.

Has Azure Front Door had more outages since October 2025?

No comparable incident has surfaced in the reporting reviewed here. Microsoft said it blocked configuration changes between October 29 and November 5, 2025 while deploying hardening fixes, and described a “shift-left” validation approach meant to catch bad configurations before they reach production, rather than after.

What exactly did Microsoft change after the October 2025 Front Door outages?

Microsoft’s Azure Networking team said it is rebuilding configuration validation, propagation safeguards and runtime behavior so no sequence of changes can destabilize the data plane. The October post is the first of a planned series; Microsoft said later posts would cover tenant isolation and recovery improvements specifically.

Why didn’t Azure’s regional failover stop the October 29 outage from going global?

Because Front Door’s configuration protection and propagation system runs as one global control plane, not one per region. Once incompatible metadata got past that layer, it reached edge sites on every continent at roughly the same time, the identical structural weakness behind AWS’s Frankfurt-triggered CloudFront failure nine months later.

Continue Reading
Click to comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Trending