Upstash unavailability in FRA region
Final update
## Impact From about 12:00 to 13:05 UTC, some clients got connection timeouts to their Redis databases. Affected were databases hosted in fra or gig, and clients whose connections were routed through fra or gig, even if their database is hosted in another region. ## What happened During a planned rolling upgrade, replicas in fra and gig did not finish draining and were left out of service. They were returned to Fly routing before they were ready to accept connections. Connections routed to them then timed out. ## What we're changing * We will be adding tooling to return a replica to service safely after maintenance. * We will be adding safety checks to our upgrade process. * We will make the upgrade procedure more resilient to errors like this one.
Timeline
- Postmortem · Sep 24, 07:22 UTC
## Impact From about 12:00 to 13:05 UTC, some clients got connection timeouts to their Redis databases. Affected were databases hosted in fra or gig, and clients whose connections were routed through fra or gig, even if their database is hosted in another region. ## What happened During a planned rolling upgrade, replicas in fra and gig did not finish draining and were left out of service. They were returned to Fly routing before they were ready to accept connections. Connections routed to them then timed out. ## What we're changing * We will be adding tooling to return a replica to service safely after maintenance. * We will be adding safety checks to our upgrade process. * We will make the upgrade procedure more resilient to errors like this one.
- Resolved · Sep 23, 14:03 UTC
This incident has been resolved.
- Monitoring · Sep 23, 13:08 UTC
A fix has been implemented and we are monitoring the results.
- Investigating · Sep 23, 12:42 UTC
We are investigating an issue affecting Upstash services in the fra region. Customers may experience connection failures or service unavailability. We are working with Upstash to identify the cause and restore service. We’ll provide an update as soon as we have more information.
More from Fly.io
Full history| Started | Incident | Impact | Duration |
|---|---|---|---|
| Sep 2416:31 UTC | State database issues in GRU | major | 2h 40m |
| Sep 2414:51 UTC | Private Networking issues (6PN) | major | 46m |
| Sep 2318:44 UTC | Partial Sprites outage | major | 7h 28m |
| Sep 2314:47 UTC | Elevated private networking errors | major | 36m |
| Sep 2306:18 UTC | IPv6 Networking Issues in DFW | major | 5h 57m |
| Sep 1504:08 UTC | Depot builder failures | minor | 1h 25m |
Sources: vendors' own status pages, published postmortems and SEC 8-K Item 1.05 filings, read daily. Times as reported. Logos via logo.dev; trademarks belong to their owners.