Skip to content
Cloudera · Data and observabilityJan 31, 2024, 14:40 UTC 2 years ago

Degraded connectivity issues between Control plane and workload clusters

MinorNot disclosedUpdated 2h ago
Jan 31, 14:40 UTCJan 31, 10:00 UTC
Duration
0m
Impact
Minor
Root cause
Not disclosed
Cloudera, 90 days
1 incidents
Affected
Not listed by the vendor.
Status page

Final update

We recently implemented an upgrade to our internal proxy service, responsible for directing requests to workloads through the Cluster Connectivity Management \(CCM\) system. This update aimed to enhance the identification of healthy routes for traffic flow. However, an unforeseen issue impacted this functionality, causing routing through potentially unhealthy paths for selected configurations. The same was fixed by rolling back the release.  It's important to note that this problem did not appear during testing in our lower environments, where rigorous evaluation precedes production deployments. Since the incident, we have promptly implemented additional safeguards to effectively identify and prevent similar issues from affecting production systems in the future. We sincerely apologize for any inconvenience this service disruption may have caused. We remain committed to delivering a reliable and robust platform, and appreciate your understanding.

Timeline

  1. Postmortem · Feb 14, 15:20 UTC
    We recently implemented an upgrade to our internal proxy service, responsible for directing requests to workloads through the Cluster Connectivity Management \(CCM\) system. This update aimed to enhance the identification of healthy routes for traffic flow. However, an unforeseen issue impacted this functionality, causing routing through potentially unhealthy paths for selected configurations. The same was fixed by rolling back the release.  It's important to note that this problem did not appear during testing in our lower environments, where rigorous evaluation precedes production deployments. Since the incident, we have promptly implemented additional safeguards to effectively identify and prevent similar issues from affecting production systems in the future. We sincerely apologize for any inconvenience this service disruption may have caused. We remain committed to delivering a reliable and robust platform, and appreciate your understanding.
  2. Resolved · Jan 31, 10:00 UTC
    We observed degraded connectivity issues between control plane and workload for a subset of customers. This impacted operations like orchestration, auto-scaling and administration of workloads from control plane. The incident has been resolved now and no action is needed from customers. A root cause analysis (RCA) will be published within seven business days.

More from Cloudera

Full history
StartedIncidentDuration
Sep 323:18 UTC4 weeks agoCloudera Management Console not accessible in US Control Plane3h 29m
May 1118:12 UTC4 months agoIntermittent Performance Issues - US Control Plane0m
Mar 1822:03 UTC6 months agoIntermittent Management Console Access Issues Across US, EU, and AP Regions0m
Sep 2510:23 UTC1 year agoFreeIPA connectivity issues6h 38m
Sep 2419:42 UTC1 year agoIntermittent Performance and Access Issues with the Cloudera Management Console21h 19m
Aug 1314:33 UTC1 year agoDataHubs, DataLakes and FreeIPA are unreachable in US region45m

Outages by email

Saturday mornings: the week's major outages, new postmortems and disclosed breaches, only in weeks that had some.

Double opt-in. Unsubscribe any time.