Skip to main content

The September 30, 2026 Railway Outage

· 14 min read
Hrishikesh Barua
Founder, IncidentHub
IncidentHub

Introduction​

Railway-hosted domains returned HTTP 404 to new connections in all four Railway regions on September 30, 2026, for about five minutes between roughly 07:35 and 07:40 UTC. This happened after a new version of Railway's routing service went live before the database schema change it depended on had been applied. Railway's report says its running workloads were not affected and no data was lost, and existing connections stayed up, although some private networking lookups between services failed in the same window.

Railway's first public status update went out at 07:48 UTC, about eight minutes after routing had recovered, and described itself as a "post facto" report. The incident was marked resolved at 08:17 UTC, and Railway published a detailed incident report on its blog at 22:42 UTC the same day.

Railway sep-30-2026 outage

Methodology and Coverage​

This summary draws on the following sources, all fetched on October 3, 2026:

  • Railway's incident report on its blog, written by Angelo Saraceno, for the cause, the internal timeline and the remediation list.
  • Railway's status page incident report, for the two public updates and the affected components.
  • Railway's status history page, for Railway's earlier 2026 incidents.
  • Railway's regions documentation, which lists four deploy regions: US West Metal (California), US East Metal (Virginia), EU West Metal (Amsterdam) and Southeast Asia Metal (Singapore).
  • Railway's May 19, 2026 incident report, for the vendor's own recent history.
  • IncidentHub's own record of Railway's status page and user reports.

Note: The internal times in the timeline (merge, rollout, schema update, detection) come only from Railway's report.

Timeline of the Outage​

Rows without a source tag (e.g. [Railway status page]) come from Railway's blog writeup.

UTCEvent
Sep 30 07:29An engineer merged a change to Railway's routing control plane after testing it in staging. The change included a database schema update and a new version of the routing service.
~07:35The new routing service was rolled out in every region. Railway says a race in the pipeline configuration ran the schema update in parallel with the deployment. Route lookups began to fail and Railway's internal alerting system notified the on-call team.
07:37[IncidentHub] recorded the first user reports of a possible Railway outage.
~07:40Railway's CI applied the schema update, the routing service retried its lookups, and domains began routing normally again.
07:44The cause was identified - the routing service came online before the schema it depended on.
07:47Railway confirmed routing had fully recovered. No rollback of the deployment was needed.
07:48[Railway status page] Monitoring: domains hosted on Railway returned 404 errors, the issue lasted about five minutes and has been mitigated, and this is a post facto report.
08:17[Railway status page] Resolved: the global routing service returned a not-found response after "a rollout of a network config change", the system self-healed, and impact to new connections lasted five minutes. Railway says it will update the status report with a published cause.
22:42Railway published its incident report on the Railway blog.

Impact window (from Railway incident report): about 07:35 UTC to about 07:40 UTC, about 5 m, from the routing service rollout to the schema update being applied. From the first failures to Resolved on the status page was about 42 m.

Sources:

Timeline of the September 30, 2026 Railway outage against its status page updates

Impact and Root Cause​

Railway was preparing to ship an authentication layer that lets customers put a login in front of their services. According to Railway, that feature needed a modified data type in the database that stores routing data, along with a new version of the routing service that reads the new table format. The two pieces deploy through separate pipelines.

An Ordering Problem​

The incident report says that due to a race condition in the configuration, the pipelines did not follow the correct order of deployment, and both ran in parallel.

The correct order should have been:

  1. Schema update rollout
  2. Routing service deployment

Although the system recovered by itself as the new database schema completed rolling out, users saw errors until that happened.

From the report:

The schema update runs as a migration job within one system, and the routing service is built by another. When the routing service was part of a larger monolith (aka the non-distributed control plane), a single pipeline enforced that order. After we split it out, into the distributed control plane to add additional resiliency, the entire system was parallelized.

The failed lookups turned into a platform-wide 404 because of how Railway's origin proxies handle a failed lookup. Railway's report says the origin proxies ask the routing service which workload a request for a domain should go to, and treat an answer of "no routes" as a domain that does not exist. During the incident the origin proxies handled a failed lookup the same way. Railway also says the change itself had modified the origin proxies' fallback behavior, so the origin proxies could not fall back to the last route they had seen for that domain, and Railway names this as the reason for the widespread impact. The same routing service resolves private network DNS, which is why some private networking lookups failed in the same window.

Services Affected​

  • New connections to custom domains and *.up.railway.app domains, which returned HTTP 404 in all regions.
  • Some private networking lookups between services.

What kept working, on Railway's word:

  • Running workloads, which kept running throughout. Railway says no data was lost.
  • Existing connections, which were not disrupted.
  • The fix itself needed no rollback. Routing recovered on the next retry once the schema update finished.

The Fallback Change​

Railway says it ships features disabled, after extensive testing in development and staging, and this change was tested in staging before the merge. However, the part of the change that altered the origin proxies' fallback behavior was evidently live in production during the race, because Railway's own report gives it as the reason the origin proxies could not fall back to the last known route. The report does not say whether that part was meant to be gated with the rest of the feature.

The fallback change probably decided the size of this outage.

Railway's Routing - Two 404 Outages in 2026​

Railway's report says the distributed control plane was "introduced recently" so that routing keeps working during a major cloud outage. Four months earlier, on May 19, 2026, Railway lost its Google Cloud-hosted infrastructure for roughly eight hours after Google Cloud suspended its production account. In that report Railway explained that its edge proxies cached routing tables from a network control plane hosted in Google Cloud. Once the cache expired, workloads on Railway Metal and AWS that were still running began returning 404 errors because the edge could no longer resolve routes, and Railway said it was working on removing that dependency.

Railway's status history lists other edge and domain incidents earlier in 2026, among them "A subset of domains are failing with a 503 error page" on February 21, "Elevated error rates on Edge Network" on March 1 (whose resolution update describes elevated 503 rates across all regions), and "Request timeouts and Service Unavailable responses on Edge Network" on March 24 to 25. The titles alone do not show a shared cause.

IncidentHub recorded 160 Railway incidents between January 1 and September 30, 2026.

MonthCount
2026-0113
2026-0219
2026-0331
2026-0424
2026-0521
2026-0625
2026-0711
2026-0812
2026-094

The Recovery​

Railway did not roll back. The routing service kept retrying its lookups, and once the schema update was applied at about 07:40 UTC the next retry succeeded and routing resumed. Railway confirmed full recovery at 07:47 UTC.

What Railway Says Comes Next​

From Railway's report:

  • Serial ordering of migration rollouts. Railway is unifying the deployment pipeline so that deploy order and blast radius are checked before merge, as they were when the system was a monolith.
  • Schema validation before go-live. Each rollout validates the schema version it was written against, and the service refuses to go live until the matching database schema is in place.
  • Shard-based rollouts per region. Railway says it already uses shard-based rollouts for most other systems and will add them to the distributed control plane to isolate impact.

Railway says that every major outage stops work on the affected system until guardrails ship, followed by a root cause analysis and a review of operating procedures. Railway says it has tested the new procedure and confirmed that it mitigates the root cause.

What This Means for Monitoring Your Cloud Providers​

If your team runs services on Railway, this outage is a useful test of your own monitoring, because it was short, it returned an unexpected status code, and it was over before the vendor said anything. Most of the checks below apply to any platform you deploy on.

Checking Only for 5xx Errors​

A check that alerts on 5xx errors or timeouts would have stayed quiet, because Railway's origin proxies answered every new connection with a 404. Ideally, point an external probe at each of your Railway-hosted domains, expect a 200 (or the exact response you know is correct), and alert on anything else.

Watching Error Rates Inside Your Own Service​

Your internal telemetry would not have shown these 404s. The origin proxies responded with 404 on their own, so the requests never reached your service, and what your dashboards would most likely have shown was a drop in traffic. Add an alert on a sudden fall in request volume, and run at least one probe from outside Railway.

Leaving Private Networking Unchecked​

Some private networking lookups between services failed in the same window, and a probe against your public domain does not exercise that path. Add a health endpoint that makes a call to another of your services over private networking and fails if that call fails.

Alert Thresholds That Are Too High for Short Outages​

A probe that runs every five minutes and needs two failures in a row before it alerts can miss an outage of this length. However, this needs to be tuned to each service's needs. If you set a higher threshold for alerts, you might miss short outages. If you set a lower threshold, you might get false positives.

Conclusion​

Railway-hosted domains returned 404 to new connections in all four regions for about five minutes on September 30, 2026, between roughly 07:35 and 07:40 UTC, according to Railway's own report. Railway attributes the outage to a routing service build that went live before the schema migration it depended on, made possible by pipelines that ran the two in parallel, and made platform-wide by a change to how the origin proxies fall back when a lookup fails. The outage ended before Railway's status page said anything, which is the reverse of the long tail seen in longer incidents, where "resolved" appears while some customers are still down.


FAQ​

How long did the September 30, 2026 Railway outage last?

About five minutes, between roughly 07:35 and 07:40 UTC, according to Railway's incident report. The status page incident ran from the first update at 07:48 UTC to resolution at 08:17 UTC.

Was any data lost?

No, according to Railway. Its report says workloads kept running throughout and only access to them was affected.

Why did Railway domains return 404 instead of a 5xx error?

Because Railway's origin proxies handled a failed route lookup the same way they handle a domain with no routes, and return 404 for both, according to Railway. The same change had modified the origin proxies' fallback, so they could not use the last route they had seen.

Was private networking affected?

Yes, partly. Railway's report says some private networking lookups between services failed, because the same routing service resolves private network DNS. The status page listed only the public networking components.

How do I detect a Railway outage like this one?

With an external probe against your own Railway-hosted domain that alerts on any response other than the one you expect. A check that only looks for 5xx errors or timeouts would have missed this outage, because Railway's origin proxies returned 404.

Is this outage related to Railway's May 19, 2026 outage?

Not by cause. The May 19 outage followed Google Cloud's suspension of Railway's account, whereas September 30 was a deployment ordering race. Both ended with Railway's proxies returning 404 when they could not resolve routes.


IncidentHub is not affiliated with any of the services and vendors mentioned in this article. All logos and company names are trademarks or registered trademarks of their respective holders. This summary is independent and not affiliated with or endorsed by Railway or any of the services and vendors mentioned in this article.

This article was first published on the IncidentHub blog.

You might also like: