Elevated wait times for machine jobs
Final update
## Summary Since October 1, CircleCI customers have experienced multiple incidents which have caused delays and failures in customer pipeline execution. There are two \(unrelated\) causes, both of which the CircleCI engineering team is actively mitigating: * Available Cloud Computing Capacity * Internal Infrastructure ## Available Cloud Computing Capacity * **Problem**: Demand for high-performance cloud instances is rising extremely rapidly across the industry, which puts pressure on the cloud provider instance types that we use to execute customer jobs. * **Incidents**: * [Elevated wait times for machine jobs](https://status.circleci.com/incidents/jn756xtsy4xg) * [Increased task wait times for Docker Gen2](https://status.circleci.com/incidents/39rfrjmk5pvd) * [Elevated level of infra fails on customer jobs](https://status.circleci.com/incidents/0z31ldjd8m5v) * [Delay on starting Machine Job Tasks](https://status.circleci.com/incidents/tj6wdnfjwr64) * [Delays starting Gen 2 Docker Jobs](https://status.circleci.com/incidents/22mrv46g6n7g) * **Mitigations and Resolutions**: * We are expanding our set of cloud computing regions to include additional regions with available high-p
Timeline
- Postmortem · Oct 8, 17:54 UTC
## Summary Since October 1, CircleCI customers have experienced multiple incidents which have caused delays and failures in customer pipeline execution. There are two \(unrelated\) causes, both of which the CircleCI engineering team is actively mitigating: * Available Cloud Computing Capacity * Internal Infrastructure ## Available Cloud Computing Capacity * **Problem**: Demand for high-performance cloud instances is rising extremely rapidly across the industry, which puts pressure on the cloud provider instance types that we use to execute customer jobs. * **Incidents**: * [Elevated wait times for machine jobs](https://status.circleci.com/incidents/jn756xtsy4xg) * [Increased task wait times for Docker Gen2](https://status.circleci.com/incidents/39rfrjmk5pvd) * [Elevated level of infra fails on customer jobs](https://status.circleci.com/incidents/0z31ldjd8m5v) * [Delay on starting Machine Job Tasks](https://status.circleci.com/incidents/tj6wdnfjwr64) * [Delays starting Gen 2 Docker Jobs](https://status.circleci.com/incidents/22mrv46g6n7g) * **Mitigations and Resolutions**: * We are expanding our set of cloud computing regions to include additional regions with available high-performance instances. * We are also working to secure additional guaranteed capacity from our cloud computing partners in our existing regions. * We are further expanding the set of instance types we can offer to customers. * We will publish a detailed Incident Report on these incidents on Octobe
- Resolved · Oct 8, 16:35 UTC
Between 12:40 UTC and 16:00 UTC on October 8, customers using Linux machine jobs and remote Docker experienced elevated wait times. The issue has been resolved and wait times have returned to normal. We thank you for your patience while our team worked on implementing a fix.
- Monitoring · Oct 8, 16:09 UTC
Wait times for customers using Linux machine jobs and remote Docker have returned to normal. We are monitoring to confirm wait times remain stable while we continue to add capacity. We will provide another update by 16:30 UTC.
- Identified · Oct 8, 15:59 UTC
Wait times continue to decrease, but customers using Linux machine jobs and remote Docker are still experiencing delays. Wait times currently average about 90 seconds, with the longest waits up to about 6 minutes on some resource classes. We are working to add capacity as quickly as possible. We will provide another update by 16:30 UTC.
- Identified · Oct 8, 15:31 UTC
Wait times continue to decrease, but customers using Linux machine jobs and remote Docker are still experiencing delays. Wait times currently average about 6 minutes. The longest waits, up to about 20 minutes, are on the 2xlarge, arm.2xlarge and gpu.nvidia.small resource classes. We are working to add capacity as quickly as possible. We will provide another update by 16:00 UTC.
- Identified · Oct 8, 15:01 UTC
A fix has been deployed and wait times are decreasing, but customers using Linux machine jobs and Remote Docker are still experiencing delays. Wait times currently average about 11 minutes, with the longest waits exceeding 35 minutes on some resource classes. We are working to add capacity as quickly as possible. We will provide another update by 15:30 UTC.
- Identified · Oct 8, 14:37 UTC
Customers using Linux machine jobs and remote Docker are experiencing elevated wait times. Wait times have started to decrease and now average about 20 minutes, with the longest waits exceeding 40 minutes on some resource classes. Our engineers have identified the issue and are working on a fix. We will provide another update by 15:00 UTC.
- Identified · Oct 8, 14:04 UTC
Customers using Linux machine jobs are experiencing elevated wait times, averaging about 40 minutes, with the longest waits exceeding 50 minutes. Most Linux machine resource classes are affected, including medium, large, xlarge, 2xlarge and their Arm equivalents. Our engineers have identified the issue and are working on a fix. We will provide another update by 14:30 UTC.
- Identified · Oct 8, 13:34 UTC
Customers using Linux machine jobs are experiencing elevated wait times, averaging about 11 minutes, with the longest waits exceeding 30 minutes on the medium, arm.medium and arm.large resource classes. Our engineers have identified the issue and are working on a fix. We will provide another update by 14:00 UTC.
- Identified · Oct 8, 13:00 UTC
Customers may be experiencing elevated wait times for machine jobs. We are working to resolve this.