Intermittent 504 errors on the Data Catalog
Final update
On June 3rd 2025 our internal systems detected intermittent 504 errors on the Data Catalog service UI. The interruption was caused by Data Catalog service components prematurely indicating they were ready to handle requests immediately after a scheduled system restart. In reality, these components were still performing essential internal data updates, a process which took approximately 24 minutes. During this period, the service was unable to fully process user requests, leading to intermittent access issues and 504 errors. To prevent similar occurrences, we have enhanced our system's readiness checks to ensure the Data Catalog service is fully prepared to serve traffic. We apologize for any inconvenience caused by the service disruption. We are committed to providing a reliable and robust platform and truly appreciate your understanding.
Timeline
- Postmortem · Jun 16, 13:40 UTC
On June 3rd 2025 our internal systems detected intermittent 504 errors on the Data Catalog service UI. The interruption was caused by Data Catalog service components prematurely indicating they were ready to handle requests immediately after a scheduled system restart. In reality, these components were still performing essential internal data updates, a process which took approximately 24 minutes. During this period, the service was unable to fully process user requests, leading to intermittent access issues and 504 errors. To prevent similar occurrences, we have enhanced our system's readiness checks to ensure the Data Catalog service is fully prepared to serve traffic. We apologize for any inconvenience caused by the service disruption. We are committed to providing a reliable and robust platform and truly appreciate your understanding.
- Resolved · Jun 4, 00:30 UTC
A temporary interruption in access to the DataCatalog service was detected by our internal monitoring systems. This interruption has been classified as transient, indicating that it was not a prolonged or persistent outage, though it did affect accessibility to the service for a brief period of time. Our technical teams are actively investigating the root cause of this disruption to understand precisely what occurred. An update will be disseminated as soon as the investigation concludes and a definitive explanation for the disruption is established.
More from Cloudera
Full history| Started | Incident | Impact | Duration |
|---|---|---|---|
| Sep 323:18 UTC4 weeks ago | Cloudera Management Console not accessible in US Control Plane | minor | 3h 29m |
| May 1118:12 UTC4 months ago | Intermittent Performance Issues - US Control Plane | none | 0m |
| Mar 1822:03 UTC6 months ago | Intermittent Management Console Access Issues Across US, EU, and AP Regions | minor | 0m |
| Sep 2510:23 UTC1 year ago | FreeIPA connectivity issues | minor | 6h 38m |
| Sep 2419:42 UTC1 year ago | Intermittent Performance and Access Issues with the Cloudera Management Console | minor | 21h 19m |
| Aug 1314:33 UTC1 year ago | DataHubs, DataLakes and FreeIPA are unreachable in US region | minor | 45m |