Clusters on the US Control Plane in an unreachable state
Nov 10, 21:13 UTCNov 10, 22:06 UTC
Duration
52m
Impact
Critical
Root cause
Network
Cloudera, 90 days
1 incidents
Affected
Cloudera Management ConsoleCloudera Data Platform (US) - CDP Management Console
Final update
The root cause of the issue was the expiration of certificates on our Cluster Connectivity Manager \(CCM\) instances. The certificates were renewed earlier this year however did not get fully installed across all services which caused the breakdown of communication between the services. To address this, we have enhanced our monitoring and devised a comprehensive remediation plan to detect and address similar incidents ahead of time. We regret the inconvenience this may have caused and strive to ensure every measure is taken to avoid similar issues in the future.
Timeline
- Postmortem · Nov 17, 21:08 UTC
The root cause of the issue was the expiration of certificates on our Cluster Connectivity Manager \(CCM\) instances. The certificates were renewed earlier this year however did not get fully installed across all services which caused the breakdown of communication between the services. To address this, we have enhanced our monitoring and devised a comprehensive remediation plan to detect and address similar incidents ahead of time. We regret the inconvenience this may have caused and strive to ensure every measure is taken to avoid similar issues in the future.
- Resolved · Nov 10, 22:06 UTC
This incident has been resolved.
- Investigating · Nov 10, 22:04 UTC
Current Status: Our teams have successfully deployed a fix for the issue and confirmed that the issue has been resolved. If you are still experiencing issues or have any questions please raise a support case with us. A root cause analysis (RCA) will be published within 7 Business days. Customer Experience: Customers on the US control plan may observe DataLake & DataHub cluster displaying an unreachable state Incident Start time: ~20:00 UTC November 10th, 2023 Incident End time: ~21:40 UTC November 10th, 2023
- Investigating · Nov 10, 21:13 UTC
Current Status: We are currently investigating an issue with our US Control Plane. We will have another update within 60 mins. Customer Experience: Customers on the US control plan may observe DataLake & DataHub cluster displaying an unreachable state Incident Start time: ~20:00 UTC November 10th, 2023
More from Cloudera
Full history| Started | Incident | Impact | Duration |
|---|---|---|---|
| Sep 323:18 UTC4 weeks ago | Cloudera Management Console not accessible in US Control Plane | minor | 3h 29m |
| May 1118:12 UTC4 months ago | Intermittent Performance Issues - US Control Plane | none | 0m |
| Mar 1822:03 UTC6 months ago | Intermittent Management Console Access Issues Across US, EU, and AP Regions | minor | 0m |
| Sep 2510:23 UTC1 year ago | FreeIPA connectivity issues | minor | 6h 38m |
| Sep 2419:42 UTC1 year ago | Intermittent Performance and Access Issues with the Cloudera Management Console | minor | 21h 19m |
| Aug 1314:33 UTC1 year ago | DataHubs, DataLakes and FreeIPA are unreachable in US region | minor | 45m |
Also caused by network
All| Started | Vendor | Incident | Impact | Duration |
|---|---|---|---|---|
| Sep 2523:00 UTC8 days ago | INC20000237 | critical | 1h 20m | |
| Sep 2215:36 UTC11 days ago | Delayed Sends/Receives - Avalanche (AVAX) | minor | 4h 57m | |
| Sep 1908:49 UTC2 weeks ago | Delayed Sends and Receives - EGLD (MultiversX) | minor | Ongoing | |
| Sep 1901:36 UTC2 weeks ago | Elevated errors in Ashburn, VA (IAD) | none | 0m | |
| Sep 1411:59 UTC2 weeks ago | Increased wait times for macOS jobs | minor | 1h | |
| Sep 1102:09 UTC3 weeks ago | Delayed mint/burn on STMT due to chain outage | major | 17h 31m |