Sep 12026 Google Cloud Multiple products in us-central1-b are experiencing network service degradation. Teams should check service health dashboards when regional network service degradation affects cloud products. Network4h 8m Network4h 8m Read Aug 272026 GitHub Incident with Copilot AI Model Providers Engineering teams should monitor provider health when using third party AI model services. Dependency2h 8m Dependency2h 8m Read Aug 262026 GitHub Incident with Actions Teams should monitor service status when experiencing issues with automated workflow systems. Capacity2h 50m Capacity2h 50m Read Aug 262026 GitHub Disruption with some GitHub services Teams should check platform status dashboards during general service disruptions. Not disclosed59m Not disclosed59m Read Aug 202026 Google Cloud We are investigating an issue where customers may experience timeouts, service degradations, errors, and elevated latencies across multiple products in the us-west1 region. Teams should monitor regional health alerts when experiencing elevated latencies and errors across cloud products. Capacity3h 40m Capacity3h 40m Read Aug 202026 GitHub Intermittent failures creating agent tasks Teams should track agent task creation pipelines to handle intermittent failures promptly. Deployment9h 54m Deployment9h 54m Read Aug 172026 GitHub Incident with GitHub.com Engineering teams should subscribe to platform status updates to remain informed during service incidents. Config change7h 36m Config change7h 36m Read Aug 62026 GitHub Incident with Actions Teams should verify service health when experiencing disruptions with deployment actions. Bug10h 42m Bug10h 42m Read Aug 62026 GitHub Incident with Pages - Deployment Lag Teams should account for deployment lag by monitoring static site hosting service health. Not disclosed1h 19m Not disclosed1h 19m Read Jul 152026 Google Cloud Google Cloud VMware Engine (GCVE), Google Cloud NetApp Volumes, and Bare Metal Solutions (BMS) services are experiencing a service outage in europe-west4-a due to a cooling failure. Teams should design multi region redundancy to withstand datacenter outages caused by cooling failures. Power12h 28m Power12h 28m Read Jul 142026 Google Cloud Google Cloud VMware Engine (GCVE) Stretched Cluster customers are experiencing zonal outages impacting network connectivity across multiple regions. Zonal outages can impact network connectivity across multiple regions for stretched cluster configurations. Network10h 40m Network10h 40m Read Jun 252026 GitHub Degradation with Webhooks, Pull Requests and Actions Service degradation can simultaneously impact multiple integrated platform features like webhooks, pull requests, and actions. Capacity37m Capacity37m Read Jun 172026 GitHub Incident with Copilot Availability Availability issues can disrupt specific AI-powered developer tools within the broader platform ecosystem. Config change54m Config change54m Read Jun 162026 GitHub Disruption with some GitHub services Platform-wide disruptions can affect only a subset of services while others remain operational. Not disclosed30m Not disclosed30m Read Jun 102026 GitHub Authentication issues related to API requests Authentication failures can prevent successful API requests across the platform. Not disclosed1h 19m Not disclosed1h 19m Read Jun 82026 GitHub Pull Requests and Issues unavailable for signed-out users Specific user segments like signed-out users may experience unique access issues to core features like pull requests and issues. Not disclosed1h 25m Not disclosed1h 25m Read Jun 42026 GitHub Copilot Code Review Failing Automated code review features can fail independently of the main platform availability. Deployment1h 57m Deployment1h 57m Read May 282026 GitHub Disruption with OpenAI Models Reliance on external AI models can lead to service disruptions if those models experience issues. Dependency1h 40m Dependency1h 40m Read May 262026 GitHub Incident with Actions and Pages Multiple related services like actions and pages can experience concurrent incidents. Not disclosed2h 21m Not disclosed2h 21m Read May 152026 GitHub Actions is experiencing degraded availability Automated workflow services can experience degraded availability independently of other platform components. Database35m Database35m Read May 72026 GitHub CCR and CCA failing to start for PR comments Not disclosed1h 54m Not disclosed1h 54m Read May 62026 GitHub Incident with Pull Requests Not disclosed3h 39m Not disclosed3h 39m Read May 62026 GitHub Disruption with some GitHub services Config change38m Config change38m Read May 62026 GitHub Incident with Actions, we are investigating reports of degraded availability Not disclosed2h 25m Not disclosed2h 25m Read May 52026 GitHub Incident with Actions Not disclosed3h 49m Not disclosed3h 49m Read May 42026 GitHub Incident with Issues and Webhooks Not disclosed55m Not disclosed55m Read Apr 272026 GitHub Disruption with some GitHub services Not disclosed2h 14m Not disclosed2h 14m Read Apr 272026 GitHub GitHub search is degraded Capacity6h 15m Capacity6h 15m Read Apr 232026 GitHub Incident with multiple GitHub services DNS1h 18m DNS1h 18m Read Apr 222026 GitHub Disruption with Copilot chat and Copilot Coding Agent Config change3h 43m Config change3h 43m Read Apr 202026 GitHub Partial degradation for code scanning default setup and for code quality Not disclosed15h 36m Not disclosed15h 36m Read Apr 162026 GitHub Incident with Codespaces Dependency3h 22m Dependency3h 22m Read Apr 132026 GitHub Incident with Pages DNS39m DNS39m Read Apr 92026 GitHub Disruption with some GitHub services Capacity25m Capacity25m Read Apr 12026 GitHub GitHub audit logs are unavailable Network4m Network4m Read Apr 12026 GitHub Disruption with GitHub's code search Deployment8h 43m Deployment8h 43m Read Mar 242026 GitHub Teams Github Notifications App is down Not disclosed2h 52m Not disclosed2h 52m Read Mar 192026 GitHub Issues with Copilot Coding Agent Not disclosed48m Not disclosed48m Read Mar 52026 GitHub Multiple services are affected, service degradation Config change2h 55m Config change2h 55m Read Mar 32026 GitHub Incident with all GitHub services Not disclosed1h 10m Not disclosed1h 10m Read Feb 272026 Google Cloud Vertex AI Gemini API customers experienced increased error rates when accessing the global endpoint. Config change1h 58m Config change1h 58m Read Feb 122026 GitHub Disruption with some GitHub services Network34m Network34m Read Feb 122026 GitHub Incident with Codespaces Deployment2h 3m Deployment2h 3m Read Feb 92026 GitHub Incident with Issues, Actions and Git Operations Not disclosed1h 8m Not disclosed1h 8m Read Feb 92026 GitHub Copilot Policy Propagation Delays Not disclosed17h 28m Not disclosed17h 28m Read Feb 92026 GitHub Incident with Pull Requests Not disclosed1h 21m Not disclosed1h 21m Read Feb 92026 GitHub Notifications are delayed Database3h 35m Database3h 35m Read Dec 52025 Cloudflare Body parsing change for a React Server Components fix causes 25-minute outage Urgent security mitigations still need a staged rollout; speed of response is exactly when a global config push is most dangerous. Config change25m Config change25m Read Nov 182025 Asana Partial downtime for Asana's MCP server Dependency0m Dependency0m Read Nov 182025 Cloudflare Oversized Bot Management feature file breaks the core proxy Treat internally generated config files like user input: validate size and shape before they propagate, and keep global kill switches for each feature. Config change5h 46m Config change5h 46m Read Oct 202025 AWS DynamoDB DNS failure takes down US-EAST-1 Automation that manages critical DNS needs its own guard against writing an empty record, and dependent services need to recover from a stale state on their own. DNS14h 32m DNS14h 32m Read Aug 52025 Anthropic Three infrastructure bugs intermittently degrade Claude responses Quality regressions hide in normal variance; continuous evaluations on production traffic catch what user reports cannot. Bug44d Bug44d Read Jul 142025 Cloudflare 1.1.1.1 public DNS resolver unreachable for 62 minutes Latent config errors can sit for weeks; progressive deployment and legacy-system cleanup matter as much as the change that finally triggers them. Config change1h 2m Config change1h 2m Read Jun 122025 Cloudflare Workers KV storage failure cascades to Access, WARP and more A shared internal primitive is a single point of failure for everything built on it; know which products have a hard dependency and give them a fallback. Dependency2h 28m Dependency2h 28m Read Jun 122025 Google Cloud Service Control crash loop returns 503s across Google Cloud New code paths belong behind feature flags, and globally replicated policy data needs staged propagation like any binary. Config change3h Config change3h Read Feb 62025 Cloudflare R2 object storage disabled during a phishing report remediation Abuse tooling needs the same guardrails as production changes: scope checks and a second pair of eyes before an action can disable a whole service. Operator1h 22m Operator1h 22m Read Dec 112024 OpenAI New telemetry service overwhelms Kubernetes control planes Test changes at production cluster size, and keep break-glass access to the control plane that does not depend on the thing that is failing. Deployment4h 22m Deployment4h 22m Read Jul 192024 CrowdStrike Falcon Channel File 291 update crashes Windows hosts worldwide Content and configuration updates need the same staged rollout, validation and customer control as code releases. Config change1h 18m Config change1h 18m Read Mar 82023 Datadog OS update breaks networking across regions Automatic updates are deployments: stagger them, and never let the same change land on every region in the same hour. Deployment26h 55m Deployment26h 55m Read Apr 52022 Atlassian Maintenance script deletes 883 customer sites Deletion should be soft by default, and bulk restores need to be rehearsed at the scale of your largest possible mistake. Operator12d 16h Operator12d 16h Read Dec 72021 AWS Internal network congestion disrupts US-EAST-1 Retry storms turn a small change into congestion; clients need backoff, and monitoring must not share the network it monitors. Network7h 10m Network7h 10m Read Jun 82021 Fastly Customer configuration triggers latent bug, 85% of network errors Customer configuration is untrusted input to a shared fleet; isolate its blast radius and test the bug classes it can reach. Bug2h 48m Bug2h 48m Read Jan 42021 Slack Overloaded AWS Transit Gateway takes Slack down on the first workday of 2021 Managed network components have scaling limits too; pre-warm for known traffic spikes and make sure provisioning survives the incident it is meant to fix. Capacity2h 18m Capacity2h 18m Read Jul 22019 Cloudflare WAF regular expression exhausts CPU worldwide for 27 minutes Rules and regexes are code; stage them, and use an engine with guaranteed linear time for untrusted input. Deployment27m Deployment27m Read Feb 282017 AWS Mistyped command removes S3 index servers in US-EAST-1 Tools should refuse to remove capacity below a safe minimum, and the status page cannot depend on the system it reports on. Operator4h 17m Operator4h 17m Read Jan 312017 GitLab Primary database data accidentally deleted, 18-hour restore A backup is only real once a restore has been tested; destructive commands on production hosts need an unmistakable prompt. Operator19h Operator19h Read