Degraded Experience - Failing to finalize some CCA Jobs
Final update
Between 2026-01-30 19:06 UTC and 2026-01-30 20:04 UTC, Copilot Coding Agent experienced sessions getting stuck, with a mismatch between the UI-reported session status and the underlying Actions and job execution state. Impacted users could observe Actions finish successfully but the session UI continuing to show in-progress state, or sessions remaining in queued state. The issue was caused by a feature flag that resulted in events being published to a new Kafka topic. Publishing failures led to buffer/queue overflows in the shared event publishing client, preventing other critical events from being emitted. We mitigated the incident by disabling the feature flag and redeploying production pods, which resumed normal event delivery. We are working to improve safeguards and detection around event publishing failures to reduce time to mitigation for similar issues in the future.
Timeline
- Resolved · Jan 30, 21:22 UTC
Between 2026-01-30 19:06 UTC and 2026-01-30 20:04 UTC, Copilot Coding Agent experienced sessions getting stuck, with a mismatch between the UI-reported session status and the underlying Actions and job execution state. Impacted users could observe Actions finish successfully but the session UI continuing to show in-progress state, or sessions remaining in queued state. The issue was caused by a feature flag that resulted in events being published to a new Kafka topic. Publishing failures led to buffer/queue overflows in the shared event publishing client, preventing other critical events from being emitted. We mitigated the incident by disabling the feature flag and redeploying production pods, which resumed normal event delivery. We are working to improve safeguards and detection around event publishing failures to reduce time to mitigation for similar issues in the future.
More from GitHub
Full history| Started | Incident | Impact | Duration |
|---|---|---|---|
| Sep 2416:51 UTC | Disruption with billing information updates | minor | 3h 50m |
| Sep 2310:11 UTC | Incident across several services | minor | 18h 44m |
| Sep 2022:13 UTC | Incident with Pull Requests | minor | 1h 9m |
| Sep 1720:59 UTC | Elevated rate of errors for OpenAI models provided by Copilot | minor | 50m |
| Sep 1607:20 UTC | Degradation with Gemini 3.8 Flash | major | 10h 28m |
| Sep 1519:11 UTC | Disruption with some GitHub services | minor | 49m |
Also caused by configuration change
All| Started | Vendor | Incident | Impact | Duration |
|---|---|---|---|---|
| Sep 2423:36 UTC | Key Value service restarts | minor | 0m | |
| Sep 2414:08 UTC | The Salesforce integration is unavailable for some customers | none | 7h 18m | |
| Sep 2218:20 UTC | Phone Number APIs and Console Were Returning Incorrect 404 Responses | none | 0m | |
| Sep 1622:30 UTC | Hyperdrive Elevated Origin Connection Failure Rates | none | 0m | |
| Sep 1211:04 UTC | Some customers experiencing blurry image previews and download issues | minor | 8h 56m | |
| Sep 1107:18 UTC | INC20000213 | critical | 6h 1m |
Sources: vendors' own status pages, published postmortems and SEC 8-K Item 1.05 filings, read daily. Times as reported. Logos via logo.dev; trademarks belong to their owners.