Degraded connectivity issues between Control plane and workload clusters
Final update
We recently implemented an upgrade to our internal proxy service, responsible for directing requests to workloads through the Cluster Connectivity Management \(CCM\) system. This update aimed to enhance the identification of healthy routes for traffic flow. However, an unforeseen issue impacted this functionality, causing routing through potentially unhealthy paths for selected configurations. The same was fixed by rolling back the release. It's important to note that this problem did not appear during testing in our lower environments, where rigorous evaluation precedes production deployments. Since the incident, we have promptly implemented additional safeguards to effectively identify and prevent similar issues from affecting production systems in the future. We sincerely apologize for any inconvenience this service disruption may have caused. We remain committed to delivering a reliable and robust platform, and appreciate your understanding.
Timeline
- Postmortem · Feb 14, 15:20 UTC
We recently implemented an upgrade to our internal proxy service, responsible for directing requests to workloads through the Cluster Connectivity Management \(CCM\) system. This update aimed to enhance the identification of healthy routes for traffic flow. However, an unforeseen issue impacted this functionality, causing routing through potentially unhealthy paths for selected configurations. The same was fixed by rolling back the release. It's important to note that this problem did not appear during testing in our lower environments, where rigorous evaluation precedes production deployments. Since the incident, we have promptly implemented additional safeguards to effectively identify and prevent similar issues from affecting production systems in the future. We sincerely apologize for any inconvenience this service disruption may have caused. We remain committed to delivering a reliable and robust platform, and appreciate your understanding.
- Resolved · Jan 31, 10:00 UTC
We observed degraded connectivity issues between control plane and workload for a subset of customers. This impacted operations like orchestration, auto-scaling and administration of workloads from control plane. The incident has been resolved now and no action is needed from customers. A root cause analysis (RCA) will be published within seven business days.
More from Cloudera
Full history| Started | Incident | Impact | Duration |
|---|---|---|---|
| Sep 323:18 UTC4 weeks ago | Cloudera Management Console not accessible in US Control Plane | minor | 3h 29m |
| May 1118:12 UTC4 months ago | Intermittent Performance Issues - US Control Plane | none | 0m |
| Mar 1822:03 UTC6 months ago | Intermittent Management Console Access Issues Across US, EU, and AP Regions | minor | 0m |
| Sep 2510:23 UTC1 year ago | FreeIPA connectivity issues | minor | 6h 38m |
| Sep 2419:42 UTC1 year ago | Intermittent Performance and Access Issues with the Cloudera Management Console | minor | 21h 19m |
| Aug 1314:33 UTC1 year ago | DataHubs, DataLakes and FreeIPA are unreachable in US region | minor | 45m |