Skip to content
Cloudera · Data and observabilityFeb 20, 2024, 15:32 UTC 2 years ago

CML/CDE/CDF cluster upgrade

MinorBugUpdated 3h ago
Feb 20, 15:32 UTCMar 1, 19:57 UTC
Duration
10d 4h
Impact
Minor
Root cause
Bug
Cloudera, 90 days
1 incidents
Affected
Cloudera Data FlowCloudera Data EngineeringCloudera AICloudera Data Platform (US) - DataFlowCloudera Data Platform (US) - Data EngineeringCloudera Data Platform (US) - Machine LearningCloudera Data Platform (AP) - DataFlowCloudera Data Platform (AP) - Data EngineeringCloudera Data Platform (AP) - Machine LearningCloudera Data Platform (EU) - DataFlowCloudera Data Platform (EU) - Data EngineeringCloudera Data Platform (EU) - Machine Learning
Status page

Final update

On February 20, 2024, customers reported issues while performing a CML workspace upgrade. Post investigation, we identified a bug in a newly promoted build; that impacted CML/CDE/CDF upgrades; and as a result we temporarily disabled CML upgrades.  It is important to note that only the upgrade functionality was impacted, and this did not have any impact on existing workload operations.  A hotfix was deployed to production to address the bug post which CML upgrades were re-enabled.  We have implemented additional corrective measures within our automated test suite to proactively detect similar issues in the future. We sincerely apologize for any inconvenience this incident may have caused to our customers.

Timeline

  1. Postmortem · Apr 1, 18:55 UTC
    On February 20, 2024, customers reported issues while performing a CML workspace upgrade. Post investigation, we identified a bug in a newly promoted build; that impacted CML/CDE/CDF upgrades; and as a result we temporarily disabled CML upgrades.  It is important to note that only the upgrade functionality was impacted, and this did not have any impact on existing workload operations.  A hotfix was deployed to production to address the bug post which CML upgrades were re-enabled.  We have implemented additional corrective measures within our automated test suite to proactively detect similar issues in the future. We sincerely apologize for any inconvenience this incident may have caused to our customers.
  2. Resolved · Mar 1, 19:57 UTC
    Current Status: Our teams have successfully deployed a fix for the issue and confirmed that the issue has been resolved. If you are still experiencing issues or have any questions please raise a support case with us. A root cause analysis (RCA) will be published within seven business days. Customer Experience: During this time cluster upgrades for CML, CDE and CDF are impacted.
  3. Monitoring · Feb 29, 20:51 UTC
    Current Status: Our teams have identified the source of the issue and have implemented a solution which is under monitoring. Please expect further updates tomorrow Customer Experience: During this time cluster upgrades for CML, CDE and CDF are impacted
  4. Identified · Feb 29, 15:07 UTC
    Current Status: Our teams have identified the source of the issue. We are working on developing and implementing a solution to restore the service’s. We will have another update towards the end of business today. Customer Experience: During this time cluster upgrades for CML, CDE and CDF are impacted
  5. Identified · Feb 29, 08:44 UTC
    The fix is currently being validated.
  6. Investigating · Feb 20, 15:32 UTC
    We are currently working on a fix for cluster upgrade failures that have been observed in the Control Plane regions. Please hold upgrading clusters for CML, CDE and CDF in any of the Control Plane regions till further update is made available.

More from Cloudera

Full history
StartedIncidentDuration
Sep 323:18 UTC4 weeks agoCloudera Management Console not accessible in US Control Plane3h 29m
May 1118:12 UTC4 months agoIntermittent Performance Issues - US Control Plane0m
Mar 1822:03 UTC6 months agoIntermittent Management Console Access Issues Across US, EU, and AP Regions0m
Sep 2510:23 UTC1 year agoFreeIPA connectivity issues6h 38m
Sep 2419:42 UTC1 year agoIntermittent Performance and Access Issues with the Cloudera Management Console21h 19m
Aug 1314:33 UTC1 year agoDataHubs, DataLakes and FreeIPA are unreachable in US region45m

Also caused by software bug

All
StartedIncidentDuration
Sep 1518:57 UTC2 weeks agoClickPipes failing on Kinesis in AWS us-east-1ClickHouse29h 46m
Sep 318:20 UTC4 weeks agoRetroactive Incident: Twilio Personalized Support Phone Line AffectedTwilio0m
Aug 819:48 UTC8 weeks agoAlerting expressions pipeline failing when recovery settingsGrafana Labs0m
Aug 615:22 UTC8 weeks agoIncident with ActionsGitHub10h 42m
Jul 2314:14 UTC2 months ago[Medium] Issues with Box HubsBox16m
Jun 1719:00 UTC3 months agoIncident With WebhooksGitHub0m

Outages by email

Saturday mornings: the week's major outages, new postmortems and disclosed breaches, only in weeks that had some.

Double opt-in. Unsubscribe any time.