The primary database was under high CPU load from 1:11PM-1:29PM PT. The impact of this is that our website and API were partially unavailable \(more precisely, at 80-90% availability\) during that time.
### Scope
What was affected: This affected most services including the website, builds, submissions, CI/CD jobs, and publishing updates.
What wasn’t affected: Serving updates to end users maintained 100.000% availability. App performance measurements sent to Observe were also unaffected.
### Root cause analysis
An internal orchestrator service used to update the state of CI/CD jobs, including build jobs and store submission jobs, was updated to retry upon application-level failures. We are investigating more thoroughly and currently believe this led to a cascade of retries that caused more failures once the database was under more load than it could handle.
### Remediation
The commit that changed retry behavior was rolled back and we are clearing the queue of retries.
For a more robust solution, the internal orchestrator must gracefully handle backpressure from the application server and database, and to apply more backoff.
Timeline
Postmortem · Oct 5, 20:40 UTC
The primary database was under high CPU load from 1:11PM-1:29PM PT. The impact of this is that our website and API were partially unavailable \(more precisely, at 80-90% availability\) during that time.
### Scope
What was affected: This affected most services including the website, builds, submissions, CI/CD jobs, and publishing updates.
What wasn’t affected: Serving updates to end users maintained 100.000% availability. App performance measurements sent to Observe were also unaffected.
### Root cause analysis
An internal orchestrator service used to update the state of CI/CD jobs, including build jobs and store submission jobs, was updated to retry upon application-level failures. We are investigating more thoroughly and currently believe this led to a cascade of retries that caused more failures once the database was under more load than it could handle.
### Remediation
The commit that changed retry behavior was rolled back and we are clearing the queue of retries.
For a more robust solution, the internal orchestrator must gracefully handle backpressure from the application server and database, and to apply more backoff.
Resolved · Oct 5, 20:11 UTC
Due to high database load, the Expo website and several services including submitting CI/CD jobs, build jobs, and store submissions were partially unavailable. The issue has subsided as of 1:30PM PT and we are continuing to diagnose the root cause and monitor the health of the services.