Skip to main content

The July 23 2026 Azure West US Outage: IP Route Removal and Downstream Impact

· 9 min read
Hrishikesh Barua
Founder, IncidentHub
IncidentHub

Introduction

On July 23, 2026, Microsoft Azure experienced a connectivity outage in the West US region that blocked traffic entering or leaving the region for nearly five hours. Workloads that stayed entirely inside West US were not affected. Microsoft's preliminary Post Incident Review (PIR) attributes the failure to a bug in maintenance request conversion software that removed IP routes from more devices than intended during routine device maintenance.

IncidentHub detected a wave of downstream SaaS outages tied to the Azure incident. Acknowledgement times on those status pages varied widely - from a few minutes to more than three hours after Azure's impact began.

IncidentHub's Azure West US Outage 23 July 2026

Methodology and Coverage

IncidentHub monitors Microsoft Azure services via public status pages, and also tracks outages in downstream services that depend on them in real time.

Not every dependent SaaS will name Azure on its status page, so the cascade count is probably lower than the actual number. A regional networking failure that cuts ingress and egress still has a wide blast radius for any SaaS that has presence in that region.

Timeline of the Outage

UTCEvent
14:44Device maintenance started. Earliest suspected customer impact. Multiple Azure services detected degradation.
14:45Networking, service teams, and incident responders reviewed traffic anomalies, routing behavior, packet loss, and recent changes.
15:00-16:00Impact scope and blast radius validated from reports and telemetry. Large-scale WAN route churn investigated.
16:00-17:45Investigation narrowed to abnormal routing and improper route advertisement; correlated with recent fiber maintenance activity.
17:45Rollback of the maintenance change request initiated.
18:26Rollback completed. Monitoring of recovering telemetry and services began.
19:41All impacted services recovered.

Impact window: 14:44-19:41 UTC (4 h 57 m).

Network restored: 18:26 UTC. Full service recovery declared at 19:41 UTC.

Source: Microsoft Azure Preliminary PIR - Issues connecting to resources in West US

The Root Cause

A routine device maintenance activity required isolating specific network paths. Azure's process converts those requests into system-readable instructions and checks that at least one of two redundant paths remains healthy, with safety checks intended to keep the impact minimal.

A bug in the request conversion system incorrectly marked additional devices as part of the maintenance event. That removed a set of IP routes from more devices than intended - between the datacenter and the wide-area network - which disrupted traffic entering or exiting the entire West US region.

As the investigation progressed, it was discovered that the route removals originated from a datacenter in the West US region, and then the maintenance change was rolled back.

Impacted Azure and Microsoft services listed in the PIR included (and may not have been limited to) App Service, Application Gateway, Application Insights, Azure AD B2C, Azure AI Search, Azure AI Speech, Azure API Management, Azure Bastion, Azure Bot Service, Azure Cosmos DB, Azure Data Explorer, Azure Database for PostgreSQL, Azure Databricks, Azure Firewall, Azure Kubernetes Service (AKS), Azure Monitor, Azure Virtual Desktop, Azure VMware Solution, ExpressRoute Circuits, ExpressRoute Gateways, Log Analytics, Microsoft Graph, Microsoft Sentinel, Partner Center, Power BI Embedded, Virtual WAN, and VPN Gateway.

Unlike the July 16 CloudFront outage, this was a single-region networking failure rather than a global edge or control-plane issue. The pattern still rhymes with some other 2025-2026 automation and change-management failures: safety checks existed, but a bug somewhere in the workflow defeated them.

Cascade Impact on Other Services

IncidentHub detected cascading outages in 19 downstream services that acknowledged impact related to the Azure West US incident. Since not all services named their upstream dependency, the actual number is likely higher.

Two operators stood out for how they published the failure:

  • MeridianLink opened 5 separate incidents, each tied to the Azure outage and covering a different component or service in their product. This meant that they could notify their customers about the issue in a more granular way - each component separately.
  • Nintex opened 2 incidents, likewise split by component, both caused by the Azure outage.

Most other providers opened a single incident for the same upstream event. Splitting by component can make internal ownership clearer, but it can also multiply status-page noise for customers unless they are watching specific components using a status page aggregator.

A few more examples of downstream impact:

  • Datadog Integrations reported Azure metrics may be delayed because of a partial outage in the Azure Monitor API.
  • Octopus Deploy named the upstream Azure outage directly. Among the known impacted services was the Octopus.com blog, so not every customer path was equally hit.
  • MongoDB Cloud saw cluster connectivity issues for clusters hosted in Azure West US, then recovered after Azure rolled out mitigation. Issues included intra-node communication failures, nodes wrongly appearing down, and cluster modification delays.
  • Wiz reported portal and API degradation across multiple US data centers. After Azure recovered, ingestion and detection were back, but it was possible that some customers could still see delays while processing caught up.
  • Ellucian Cloud's Student Financial Success line of products went offline while Microsoft worked the West US fault, then brought it back when Azure recovered. This was an EdTech product, and we do not know the exact impact on educational institutions.
Azure West US Outage 23 July 2026 - Cascade Impact

Downstream Recovery Lag

Azure declared all its impacted services recovered at 19:41 UTC. Of the 19 downstream services IncidentHub tracked for this cascade, 16 still had open incidents well past that timestamp.

Recovery in downstream services during bigger infra outages can sometimes trigger secondary outages - from thundering herds, or from third-party rate limiting due to backed up queues getting processed. We did not observe any in incident reports in this case but that's not to say that none occurred. Also, once engineers are satisfied that an outage is mitigated, there is a period of monitoring to make sure things are absolutely ok.

The outage start times in the graph are from the status page announcements, and the huge lags from a few minutes to 3+ hours can indicate various things in the downstream services:

  • Their own detection lag
  • Announcement lag in the status page
  • Actually delayed impact

Surviving Regional Azure Outages

You cannot eliminate every dependency risk, but regional networking failures will make you think about your dependencies a bit more:

  • Know which of your workloads and SaaS vendors pin critical paths to a single Azure region, especially if that region hosts your identity management SaaS or a vendor's control plane.
  • Architect multi-region designs for mission-critical paths.
  • Treat ingress/egress cloud components as first-class dependencies. An ingress/egress failure looks like "the app is down" even when compute inside the region is fine.
  • Monitor upstream cloud status and your SaaS dependencies together using a status page aggregator like IncidentHub.
  • Expect downstream status pages to lag. Plan your customer communications based on your own monitoring.

Conclusion

The July 23 outage was nearly five hours long, confined to West US ingress and egress, and caused by a maintenance automation bug that removed more IP routes than intended. Azure restored the network by 18:26 UTC and declared full service recovery at 19:41 UTC. Many downstream SaaS vendors felt it for a longer period: incidents in 16 of the 19 dependent services IncidentHub tracked stayed open past Azure's recovery time, and acknowledgement lagged by minutes to hours.

If your stack or your vendors concentrate critical traffic in one region, a single regional failure in a critical component like networking can effectively cut off your application from your users.



FAQ

What caused the July 23, 2026 Azure West US outage?

A bug in Azure's maintenance request conversion system incorrectly marked additional devices as part of a routine maintenance event and removed IP routes from more devices than intended, cutting connectivity between a West US datacenter and the wide-area network.

How long did the Azure outage last?

Customer impact ran from 14:44 UTC to 19:41 UTC on July 23, 2026 - 4 hours and 57 minutes. The network rollback completed at 18:26 UTC; all impacted services were declared recovered at 19:41 UTC.

Was traffic inside West US affected?

No. Microsoft's PIR states that impact was limited to network traffic entering or exiting the West US region. Traffic that remained entirely within the region was not affected.

How many downstream services were affected?

IncidentHub detected cascading outages in 19 downstream services. 16 of those 19 still had open incidents after Azure's 19:41 UTC recovery.

Where is Microsoft's official report?

Microsoft published a Preliminary PIR on the Azure status history page: Issues connecting to resources in West US (tracking ID ZJV6-SGG).


IncidentHub is not affiliated with any of the services and vendors mentioned in this article. All logos and company names are trademarks or registered trademarks of their respective holders. This summary is independent and not affiliated with or endorsed by Microsoft, Azure, or any of the services and vendors mentioned in this article.

This article was first published on the IncidentHub blog.