Skip to content
GitHub · Developer toolsJan 30, 2026, 20:59 UTC

Degraded Experience - Failing to finalize some CCA Jobs

MinorConfig changeUpdated 37h ago
Jan 30, 20:59 UTCJan 30, 21:22 UTC
Duration
23m
Impact
Minor
Root cause
Config change
GitHub, 90 days
69 incidents
Affected
Not listed by the vendor.
Status page

Final update

Between 2026-01-30 19:06 UTC and 2026-01-30 20:04 UTC, Copilot Coding Agent experienced sessions getting stuck, with a mismatch between the UI-reported session status and the underlying Actions and job execution state. Impacted users could observe Actions finish successfully but the session UI continuing to show in-progress state, or sessions remaining in queued state. The issue was caused by a feature flag that resulted in events being published to a new Kafka topic. Publishing failures led to buffer/queue overflows in the shared event publishing client, preventing other critical events from being emitted. We mitigated the incident by disabling the feature flag and redeploying production pods, which resumed normal event delivery. We are working to improve safeguards and detection around event publishing failures to reduce time to mitigation for similar issues in the future.

Timeline

  1. Resolved · Jan 30, 21:22 UTC
    Between 2026-01-30 19:06 UTC and 2026-01-30 20:04 UTC, Copilot Coding Agent experienced sessions getting stuck, with a mismatch between the UI-reported session status and the underlying Actions and job execution state. Impacted users could observe Actions finish successfully but the session UI continuing to show in-progress state, or sessions remaining in queued state. The issue was caused by a feature flag that resulted in events being published to a new Kafka topic. Publishing failures led to buffer/queue overflows in the shared event publishing client, preventing other critical events from being emitted. We mitigated the incident by disabling the feature flag and redeploying production pods, which resumed normal event delivery. We are working to improve safeguards and detection around event publishing failures to reduce time to mitigation for similar issues in the future.

More from GitHub

Full history
StartedIncidentDuration
Sep 2416:51 UTCDisruption with billing information updates3h 50m
Sep 2310:11 UTCIncident across several services18h 44m
Sep 2022:13 UTCIncident with Pull Requests1h 9m
Sep 1720:59 UTCElevated rate of errors for OpenAI models provided by Copilot50m
Sep 1607:20 UTCDegradation with Gemini 3.8 Flash10h 28m
Sep 1519:11 UTCDisruption with some GitHub services49m

Also caused by configuration change

All
StartedIncidentDuration
Sep 2423:36 UTCKey Value service restartsRender0m
Sep 2414:08 UTCThe Salesforce integration is unavailable for some customersHubSpot7h 18m
Sep 2218:20 UTCPhone Number APIs and Console Were Returning Incorrect 404 ResponsesTwilio0m
Sep 1622:30 UTCHyperdrive Elevated Origin Connection Failure RatesCloudflare0m
Sep 1211:04 UTCSome customers experiencing blurry image previews and download issuesSlack8h 56m
Sep 1107:18 UTCINC20000213Snowflake6h 1m

Sources: vendors' own status pages, published postmortems and SEC 8-K Item 1.05 filings, read daily. Times as reported. Logos via logo.dev; trademarks belong to their owners.

Outages by email

Saturday mornings: the week's major outages, new postmortems and disclosed breaches, only in weeks that had some.

Double opt-in. Unsubscribe any time.