Copilot Coding Agent failing to start some jobs
Final update
Between 15:20 and 20:18 UTC on Thursday April 2, Copilot Cloud Agent entered a period of reduced performance. Due to an internal feature being developed for Copilot Code Review, the Copilot Cloud Agent infrastructure started to receive an increased number of jobs. This load eventually caused us to hit an internal rate limit, causing all work to suspend for an hour. During this hour, some new jobs would time out, while others would resume once rate limiting ended. Roughly 40% of jobs in this period were affected. Once the cause of this rate limiting was identified, we were able to disable the new CCR feature via a feature flag. Once the jobs that were already in the queue were able to clear, we didn't see additional instances of rate limiting afterwards. This was the same incident declared in https://www.githubstatus.com/incidents/d96l71t3h63k
Timeline
- Resolved · Apr 2, 16:30 UTC
Between 15:20 and 20:18 UTC on Thursday April 2, Copilot Cloud Agent entered a period of reduced performance. Due to an internal feature being developed for Copilot Code Review, the Copilot Cloud Agent infrastructure started to receive an increased number of jobs. This load eventually caused us to hit an internal rate limit, causing all work to suspend for an hour. During this hour, some new jobs would time out, while others would resume once rate limiting ended. Roughly 40% of jobs in this period were affected. Once the cause of this rate limiting was identified, we were able to disable the new CCR feature via a feature flag. Once the jobs that were already in the queue were able to clear, we didn't see additional instances of rate limiting afterwards. This was the same incident declared in https://www.githubstatus.com/incidents/d96l71t3h63k
More from GitHub
Full history| Started | Incident | Impact | Duration |
|---|---|---|---|
| Sep 2416:51 UTC | Disruption with billing information updates | minor | 3h 50m |
| Sep 2310:11 UTC | Incident across several services | minor | 18h 44m |
| Sep 2022:13 UTC | Incident with Pull Requests | minor | 1h 9m |
| Sep 1720:59 UTC | Elevated rate of errors for OpenAI models provided by Copilot | minor | 50m |
| Sep 1607:20 UTC | Degradation with Gemini 3.8 Flash | major | 10h 28m |
| Sep 1519:11 UTC | Disruption with some GitHub services | minor | 49m |
Also caused by configuration change
All| Started | Vendor | Incident | Impact | Duration |
|---|---|---|---|---|
| Sep 2423:36 UTC | Key Value service restarts | minor | 0m | |
| Sep 2414:08 UTC | The Salesforce integration is unavailable for some customers | none | 7h 18m | |
| Sep 2218:20 UTC | Phone Number APIs and Console Were Returning Incorrect 404 Responses | none | 0m | |
| Sep 1622:30 UTC | Hyperdrive Elevated Origin Connection Failure Rates | none | 0m | |
| Sep 1211:04 UTC | Some customers experiencing blurry image previews and download issues | minor | 8h 56m | |
| Sep 1107:18 UTC | INC20000213 | critical | 6h 1m |
Sources: vendors' own status pages, published postmortems and SEC 8-K Item 1.05 filings, read daily. Times as reported. Logos via logo.dev; trademarks belong to their owners.