{
"vendor": "Cloudera",
"slug": "cloudera",
"platform": "statuspage",
"status_url": "https://status.cloudera.com",
"last_checked": "2026-09-16T12:28:20Z",
"last_state": "ok",
"history_backfilled": true,
"first_watched": "2026-09-04T07:06:16Z",
"incidents": [
{
"body": "Current Status: Our teams have successfully implemented a solution to resolve the issue. Our teams are currently monitoring the solution and will provide another update once confirmed.\n\nCustomer Experience: Customer may have observed an Error message: \"no healthy upstream\" while accessing Cloudera Management Console.\nIf you are still experiencing issues or have any questions, please raise a support case with us. A Root Cause Analysis (RCA) will be published within seven business days.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-09-04T02:47:17.907Z",
"resolved_inferred": false,
"started_at": "2026-09-03T23:18:07.714Z",
"state": "resolved",
"title": "Cloudera Management Console not accessible in US Control Plane",
"updated_at": "2026-09-04T02:47:17.933Z",
"url": "https://stspg.io/jzl2h2lmr01s"
},
{
"body": "On March 10, 2026, between 15:23 UTC to 17:09 UTC, users may have experienced intermittent connectivity issues accessing the workloads in the Cloudera Management Console of the US region.\n\nThis disruption was traced to the termination of jumpgate-proxy pods, which entered an OOMKilled status after memory utilization exceeded configured limits. This, in turn, impacted connectivity between the control plane and the workloads within the Cloudera Management Console.\n\nService consistency was restored by increasing the memory limits and replica counts for the jumpgate-proxy deployment.\n\nTo prevent recurrence, we will implement alerting mechanisms for jumpgate pod memory utilization and pod states \\(such as CrashLoopBackoff or OOMKilled events\\). This will enable us to proactively detect and mitigate memory pressure before it impacts service availability. We will implement autoscaling to autoheal and respond to increased load.\n\nWe sincerely apologise for any inconvenience this service disruption may have caused. We appreciate your patience as we worked to restore service functionality.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-05-11T10:00:00.000Z",
"resolved_inferred": false,
"started_at": "2026-05-11T10:00:00.000Z",
"state": "postmortem",
"title": "Intermittent Performance Issues - US Control Plane",
"updated_at": "2026-05-22T19:35:43.119Z",
"url": "https://stspg.io/6t0hy3kcj8g0"
},
{
"body": "On March 10, 2026, around 19:56 UTC users may have started experiencing intermittent timeouts accessing the Cloudera Management Console across the US, EU, and AP regions.\u00a0\n\nThis disruption was traced to an internal service certificate expiration where, despite a successful renewal initiation, a misalignment in the deployment sequence prevented the updated credentials from propagating to the active service mesh. The impact was strictly limited to administrative console and API access; all existing data workloads and running environments continued to operate without interruption.\u00a0\n\nService consistency was restored by regenerating and manually applying the certificates across all affected clusters. To prevent recurrence, we have audited all regional clusters and are transitioning to an automated lifecycle for mesh-level certificates.\u00a0\n\nWe sincerely apologize for any inconvenience this service disruption may have caused. We are addressing monitoring gaps to detect these issues before they impact you, our customers. We appreciate your patience as we worked to restore service functionality.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-03-10T20:00:00.000Z",
"resolved_inferred": false,
"started_at": "2026-03-10T20:00:00.000Z",
"state": "postmortem",
"title": "Intermittent Management Console Access Issues Across US, EU, and AP Regions",
"updated_at": "2026-03-18T22:08:38.400Z",
"url": "https://stspg.io/2f2l4x28khkx"
},
{
"body": "The intermittent performance and access issues with the Cloudera Management Console and FreeIPA were triggered by a scheduled service upgrade intended to improve platform stability.\n\nOur investigation determined that a change in a core component introduced a latent configuration issue. This specific condition was not exposed in our testing environments, preventing dependent services from applying dynamic configuration updates in production and leading to the outage.\n\nWe've taken immediate action to prevent this issue from recurring:  \nSystem Fix: The problematic component change was rolled back, and a permanent patch was deployed to restore proper dynamic configuration functionality.  \nProcess Overhaul: We've implemented a more rigorous upgrade process with mandatory, near-production scale validation steps that specifically test for these types of configuration failures.  \nEnhanced Monitoring: We significantly improved our monitoring and alerting capabilities to detect these abnormal service behaviours much earlier, ensuring a faster response time.\n\nWe are dedicated to providing a reliable platform and will continue to invest in our infrastructure and processes. Thank you again for your patience and understanding",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-09-25T17:02:03.359Z",
"resolved_inferred": false,
"started_at": "2025-09-25T10:23:34.323Z",
"state": "postmortem",
"title": "FreeIPA connectivity issues",
"updated_at": "2025-10-13T20:05:00.018Z",
"url": "https://stspg.io/5ts1ff98nf42"
},
{
"body": "The intermittent performance and access issues with the Cloudera Management Console and FreeIPA were triggered by a scheduled service upgrade intended to improve platform stability.\n\nOur investigation determined that a change in a core component introduced a latent configuration issue. This specific condition was not exposed in our testing environments, preventing dependent services from applying dynamic configuration updates in production and leading to the outage.\n\nWe've taken immediate action to prevent this issue from recurring:  \nSystem Fix: The problematic component change was rolled back, and a permanent patch was deployed to restore proper dynamic configuration functionality.  \nProcess Overhaul: We've implemented a more rigorous upgrade process with mandatory, near-production scale validation steps that specifically test for these types of configuration failures.  \nEnhanced Monitoring: We significantly improved our monitoring and alerting capabilities to detect these abnormal service behaviours much earlier, ensuring a faster response time.\n\nWe are dedicated to providing a reliable platform and will continue to invest in our infrastructure and processes. Thank you again for your patience and understanding",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-09-25T17:01:35.247Z",
"resolved_inferred": false,
"started_at": "2025-09-24T19:42:22.945Z",
"state": "postmortem",
"title": "Intermittent Performance and Access Issues with the Cloudera Management Console",
"updated_at": "2025-10-13T20:04:17.193Z",
"url": "https://stspg.io/sk4jg5x4pjxb"
},
{
"body": "A recent system update, intended to enhance stability of our platform, inadvertently led to an unforeseen memory issue affecting a critical internal service responsible for platform access management.\n\n\u200c\n\nA subsequent incoming request spike caused the memory limits exhaustion, which resulted in the service to become intermittently unavailable.\n\n\u200c\n\nThe memory allocation for the service was manually increased, which restored normal operations. This change was then made permanent to circumvent any recurrence.\n\n\u200c\n\nWe have implemented more robust monitoring and alerting to detect similar issues in the future before they can impact our customers. This includes updating the existing alerts for resource exhaustion and adding new alerts based on the incident, to ensure that the right teams are notified immediately.\n\n\u200c\n\nWe are dedicated to providing a reliable and performant platform. We will continue to invest in improving our infrastructure and processes to prevent future disruptions. We appreciate your patience and understanding as we worked to resolve this issue.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-08-13T15:18:43.292Z",
"resolved_inferred": false,
"started_at": "2025-08-13T14:33:47.920Z",
"state": "postmortem",
"title": "DataHubs, DataLakes and FreeIPA are unreachable in US region",
"updated_at": "2025-08-22T18:16:24.076Z",
"url": "https://stspg.io/kjv7xw15yn3n"
}
]
}