The July 24, 2026 AWS us-west-2 Outage: Network Routing and a Long Recovery Tail
Introduction
On July 24, 2026, AWS lost connectivity between the us-west-2 (Oregon) region and the Seattle Metro. The initial impact window was 20 minutes for most and 1 hour 17 minutes for a few customers using AWS Direct Connect through EqSe2, Westin Building Exchange, Seattle. Any traffic that both started and ended inside the region kept working, whereas anything crossing the region boundary saw timeouts and errors. This included the AWS Management Console for some customers.
Downstream services took longer to recover. IncidentHub recorded 9 confirmed cascade incidents across 7 providers, and 3 possible cascade events at 3 more.
- Introduction
- Methodology and Coverage
- Timeline of the Outage
- The Root Cause
- Cascade Impact on Other Services
- The Recovery Tail
- Three Infrastructure Failures in Nine Days
- Surviving Regional Network Outages
- Conclusion
- FAQ
Methodology and Coverage
IncidentHub monitors all AWS services and regions using the AWS Health Dashboard, and separately tracks outages in downstream services that depend on them in real time.
Downstream impact is based on two factors:
- If the downstream service directly mentions AWS or Amazon Web Services in its own incident update, then it is a confirmed cascade event.
- If the downstream service does not mention AWS directly, but the time window coincides and the update names the same region, an upstream provider, or both, then we mark it as a possible cascade event.
Not all downstream providers explicitly mention AWS in their own incident updates, so the number of impacted services is likely higher than we see here.
The resolved timestamps below are also status page state, not customer impact. A provider resolving an incident at 21:28 UTC does not mean every customer was broken until 21:28 UTC, and a provider closing at 11:12 UTC does not prove nobody was.
Timeline of the Outage
| UTC | Event |
|---|---|
| 10:55 | Loss of connectivity to us-west-2 began, across several AWS services. Traffic staying inside the region kept working. |
| 11:01 | AWS engineers were automatically engaged by alerting systems. |
| 11:15 | Initial mitigation was completed, and connectivity began coming back. |
| 11:40 | First public post on the Health Dashboard, confirming an investigation into us-west-2 connectivity across several services. |
| 11:47-11:59 | A "reconvergence" event resulted in some customers experiencing intermittent connectivity as networking routes were re-established. |
| 11:59 | Routing fully back, with AWS service metrics at their pre-outage levels. |
| 12:12 | Routing restored for the Direct Connect path through EqSe2, Westin Building Exchange, Seattle. |
| 12:30 | Public update naming the root cause and confirming that mitigation work had finished. |
| 13:01 | Closing summary posted. |
Initial impact window: 10:55-11:15 UTC (20 minutes), plus route reconvergence from 11:47 to 11:59 UTC.
Direct Connect at EqSe2, Westin Building Exchange, Seattle: 10:55-12:12 UTC (1 h 17 m). AWS noted that anyone holding a redundant path through a different Direct Connect location was not impacted by the event.
Note: AWS's first public update went up at 11:40 UTC, which is 45 minutes after impact began and 25 minutes after their own initial recovery. Engineers were pulled in automatically at 11:01 UTC, so internal detection was quick.
Source: AWS Health Dashboard, us-west-2 operational issue, July 24, 2026
The Root Cause
AWS attributed the fault to the networking hardware that carries routing between the region and the Seattle Metro. Several mitigation paths were worked in parallel, and the first of them brought connectivity back at 11:15 UTC.
As the network settled after the mitigation, routes reconverged between 11:47 and 11:59 UTC, and connectivity flapped again for some customers while the new paths settled. The initial fix had worked, and the act of applying it produced a second, smaller window of impact.
Reconvergence in Network Routing
Convergence in network routing is when a set of routers agree among themselves on the network paths to be taken while routing traffic. It reflects the state of the network. Routers exchange path information among themselves constantly - so whenever there is a route change, or a new route is announced, this information takes time to propagate through the network.
In this case, as new routes were announced, and routers started to update their routing tables, there would be intermittent "dead" or slow routes until all of them agreed on the new state of the network. This reconvergence period is when customers saw intermittent connectivity issues.
AWS listed 10 affected services: AWS Direct Connect, AWS Global Accelerator, AWS Internet Connectivity, AWS IoT Core, AWS Site-to-Site VPN, Amazon API Gateway, Amazon Elastic Compute Cloud, Amazon Elastic Container Service, Amazon Elastic Load Balancing, and Amazon Virtual Private Cloud. That list is almost entirely the ingress and egress layer. Compute inside the region kept running as usual.
Cascade Impact on Other Services
Confirmed Cascade Events
Nine incidents at seven providers named AWS or Amazon Web Services in their own updates. The last column measures from 11:59 UTC, when AWS said routing was fully back.
| Provider | Opened (UTC) | Resolved (UTC) | After AWS route restore |
|---|---|---|---|
| StatusHub | 10:59 | 14:10 | 2 h 11 m |
| Cube Cloud | 11:03 | 11:12 | closed before restore |
| StatusHub | 11:04 | 14:10 | 2 h 11 m |
| StatusHub | 11:04 | 14:10 | 2 h 11 m |
| Bambu Lab | 11:12 | 13:18 | 1 h 19 m |
| SparkPost | 11:14 | 18:44 | 6 h 45 m |
| FortiSASE | 11:20 | 13:31 | 1 h 32 m |
| NinjaOne | 11:38 | 21:28 | 9 h 29 m |
| SendGrid | 12:06 | 19:31 | 7 h 32 m |
StatusHub appears three times because they opened three separate incidents, all within five minutes of each other and all closed together.
Three of these cases are worth reading carefully:
- StatusHub opened three separate incidents and kept all three open until 14:10 UTC. Their updates through that stretch said while the platform itself was healthy, the incidents were being held open purely because the upstream AWS incident had not been closed yet. It is a reminder that status page state and service state are not the same thing.
- FortiSASE first filed the incident against AWS in EMEA, where Endpoint Management and Portal access were degraded for some customers. A later update noted that AWS had posted about US West (Oregon) but that the link to what they were seeing was still unconfirmed, and pointed at a spike on DownDetector as a secondary signal. They retitled the incident to cover Oregon and EMEA once the connection was confirmed.
- SendGrid is the counter-example. DNS health checks fired at around 11:09 UTC, and by 11:21 UTC their engineers had pulled the degraded region out of load balancing. Their update said customers should see no impact and that mail sending was unaffected.
Bambu Lab, whose cloud services sit behind consumer 3D printers, said they were shifting traffic to healthy regions where they could.
Possible Cascade Events
Three more incidents at three providers line up with the outage on timing and either a region match or an unnamed upstream mention, but none of them named AWS. They are listed separately because the link is inferred rather than explicitly stated.
| Provider | Opened (UTC) | Resolved (UTC) | After AWS route restore |
|---|---|---|---|
| TecAlliance TecDoc | 11:06 | 11:36 | closed before restore |
| Render | 11:16 | 13:30 | 1 h 31 m |
| RubyGems | 11:29 | 12:06 | 7 m |
- TecAlliance TecDoc reported network issues for region-specific requests sent to US-West-2. The region matches but AWS is not named.
- Render posted a service disruption in Oregon and blamed an upstream provider without saying which one.
- RubyGems traced degraded service to an upstream provider during the same window, with no region named.
The Recovery Tail
In the Azure West US outage the previous day, IncidentHub noted that recovery can trigger secondary problems and that none had been called out in the published incident reports. This time one provider wrote it up plainly.
NinjaOne posted their account at 15:38 UTC, well after AWS had finished. Their agents had tried to reach the backend during the outage, failed, and dropped into the backoff and retry path that exists specifically to stop an entire fleet hammering a recovering service at once. Reconnection was therefore slow and randomized by design. Separately, the service that processes device status could not keep up with the volume of agents changing state, and had to be scaled out before its metrics looked healthy again. They reported more than 150,000 devices coming back in the preceding hour, with the rest still working their way through.
SparkPost shows the same shape in a different layer. Their 17:32 UTC update described extra tuning to absorb the surge of mail that had piled up behind the AWS network fault, with full delivery expected to need another 80 to 90 minutes on top. Their earlier updates describe it: outbound delivery from the US West region was blocked by 11:53 UTC while they carried on accepting and queuing messages internally, delivery restarted at 12:08 UTC, and they were still working through the backlog at 13:16 UTC.
These are the aspects that are not evident in the upstream provider's impact reports - they are often felt by customers of downstream providers.
Three Infrastructure Failures in Nine Days
If we set this alongside the two that came before it:
- July 16, 2026 - AWS CloudFront. Something in the fleet handling connections to private VPC origins stopped it loading updated configuration correctly. It was global in reach, roughly 3 h 33 m.
- July 23, 2026 - Microsoft Azure West US. A defect in how maintenance requests were translated pulled routes off a wider set of devices than the change had scoped, cutting off the datacenter from the wide-area network. Ingress and egress only, 4 h 57 m.
- July 24, 2026 - AWS us-west-2. Networking hardware on the path between the region and the Seattle Metro. Ingress and egress only, 20 minutes plus reconvergence impact.
Two of the three events broke at the same boundary - between a region and the wider network. In both cases compute, storage and databases inside the region were healthy but unreachable, which for a customer is indistinguishable from being down. AZ redundancy does not help in such cases.
Another pattern is invisible unless the downstream impact is checked - upstream duration cannot predict downstream duration. CloudFront was the longest of the three upstream events and the July 24 AWS event was by far the shortest, yet the July 24 cascade produced a long downstream tail.
Surviving Regional Network Outages
- Know which of your vendors run in us-west-2, and which of them can fail over out of it. SendGrid could do it and activated it. Bambu Lab could do it but not completely. Most of the others did not mention failover at all.
- Plan for the secondary effects like reconvergence. AWS's own timeline has a second impact window after the fix, and your monitoring should not treat the first recovery signal as the end. However, this is difficult to plan for specifically and it's far better to monitor things yourself after your upstream provider has declared recovery. This is probably one of the reasons why downstream providers take longer to declare recovery in such cases.
- Plan for thundering herd issues before it happens. NinjaOne's agents had backoff and jitter and it still took hours.
- Audit your internal queue capacities and processing rates, and put knobs in place to either throttle or scale up capacity to avoid huge backlogs becoming bottlenecks.
- Monitor upstream cloud status and your SaaS dependencies in one place.
Conclusion
The upstream fault at AWS was relatively small and quickly handled. Networking hardware between us-west-2 and the Seattle Metro failed at 10:55 UTC, AWS engineers were on it by 11:01 UTC, connectivity started returning at 11:15 UTC, and routing was fully back by 11:59 UTC. Traffic inside the region was unaffected.
Since a networking issue effectively means closing the gate, downstream providers have to wait for the gate to be opened again, or failover to other available regions, before they can start processing traffic. There were 9 confirmed downstream incidents across seven providers, plus three possible ones. None opened before AWS impact began at 10:55 UTC; 11 of the 12 were open before AWS's first public update at 11:40 UTC. The recovery tail lasted till 21:28 UTC on a short outage.
FAQ
What caused the July 24, 2026 AWS us-west-2 outage?
AWS attributed it to the networking hardware carrying routes between the us-west-2 region and the Seattle Metro. As of this writing, AWS has not published details of what was wrong with that hardware.
How long did the AWS outage last?
The initial impact ran from 10:55 to 11:15 UTC, which is 20 minutes. Routes then reconverged between 11:47 and 11:59 UTC, causing further intermittent connectivity, and routing was fully back by 11:59 UTC. Customers using AWS Direct Connect through EqSe2, Westin Building Exchange, Seattle were impacted until 12:12 UTC.
Was traffic inside us-west-2 affected?
No. AWS said traffic that stayed inside the region kept working. The impact was on traffic crossing the region boundary, which also meant some customers could not reach the AWS Management Console.
Which downstream services were affected?
IncidentHub recorded 9 confirmed incidents across 7 providers that named AWS in their own incident updates: StatusHub, Cube Cloud, Bambu Lab, SparkPost, FortiSASE, NinjaOne, and SendGrid. A further 3 incidents are counted as possible rather than confirmed: TecAlliance TecDoc reported network issues for US-West-2 requests without naming AWS, Render reported an Oregon disruption blamed on an unnamed upstream provider, and RubyGems traced degradation to an unnamed upstream provider in the same time window. The real number is likely higher.
Why did downstream services take so long to recover?
Backlog and reconnection load, mostly. NinjaOne's agents reconnected slowly by design, using backoff and jitter to avoid a thundering herd, and their device status processing had to be scaled out to absorb the volume. SparkPost had to drain a mail backlog that built up during the outage. Both incidents closed hours after AWS restored routing.
IncidentHub is not affiliated with any of the services and vendors mentioned in this article. All logos and company names are trademarks or registered trademarks of their respective holders. This summary is independent and not affiliated with or endorsed by AWS or any of the services and vendors mentioned in this article.
This article was first published on the IncidentHub blog.
You might also like:

