{
"vendor": "Ionos Cloud",
"slug": "ionos-cloud",
"platform": "statuspage",
"status_url": "https://status.ionos.cloud",
"last_checked": "2026-09-16T12:28:20Z",
"last_state": "ok",
"history_backfilled": true,
"first_watched": "2026-09-04T07:06:16Z",
"incidents": [
{
"body": "DBaaS service has recovered, as well. We are setting the incident into Monitoring status, while the teams are ensuring that all affected services operate normally.",
"first_seen": "2026-09-09T12:30:34Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"started_at": "2026-09-08T23:43:26.379Z",
"state": "monitoring",
"title": "TXL - Network Connectivity",
"updated_at": "2026-09-09T02:19:43.893Z",
"url": "https://stspg.io/yhcpwhz191tz"
},
{
"body": "We have people available now, Support is reachable as usual.",
"first_seen": "2026-09-09T12:30:34Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-09-08T21:57:50.544Z",
"resolved_inferred": false,
"started_at": "2026-09-08T14:15:48.715Z",
"state": "resolved",
"title": "Cloud Support: Limited Phone Support Availability",
"updated_at": "2026-09-08T21:57:50.572Z",
"url": "https://stspg.io/w55tst6ln1tx"
},
{
"body": "We managed to fully re-establish our service, Support is back.\nthank you for your patience.",
"first_seen": "2026-09-04T17:10:49Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-09-08T06:31:22.093Z",
"resolved_inferred": false,
"started_at": "2026-09-04T16:53:01.153Z",
"state": "resolved",
"title": "Cloud Support: Limited Phone Support Availability",
"updated_at": "2026-09-08T06:31:22.111Z",
"url": "https://stspg.io/lj94bvfrkk8y"
},
{
"body": "This incident has been resolved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-09-01T16:12:50.542Z",
"resolved_inferred": false,
"started_at": "2026-09-01T13:36:02.407Z",
"state": "resolved",
"title": "Partner Subcontracts can not access new location de/fra/1",
"updated_at": "2026-09-01T16:12:50.565Z",
"url": "https://stspg.io/wnv06cj6dsm0"
},
{
"body": "# **Preliminary Root Cause Analysis**\n\nThis Root Cause Analysis is preliminary as research is still being conducted to determine the technical root cause of the incident.\n\n## **What happened?**\n\nOn August 31, 2026, between 15:00 UTC and 15:43 ITC, and again between 16:32 UTC and 16:35 UTC, customers were unable to reach IONOS Cloud's Identity and Access Management \\(IAM\\) service and the Data Center Designer \\(DCD\\), which is dependent on this service. The disruption affected direct logins, partner and reseller portal access, and associated management consoles. The total customer-facing impact lasted approximately 48 minutes across both intervals.\n\nCloud APIs \\([api.ionos.com](http://api.ionos.com)\\) remained fully operational throughout the incident. The underlying application services were healthy at all times - the failure was confined to the network edge layer.\n\n## **How was this possible? \\(Root Cause\\)**\n\nDuring a scheduled maintenance window on August 31, 2026, a network configuration update was applied to the edge network infrastructure with the intent of optimizing routing filters. The configuration was verified as correct prior to and during application. Following the rollout, a routing propagation anomaly emerged on one edge network switch, causing asymmetric routing behavior: incoming TCP connection requests from clients were silently dropped \\(black-holed\\) at the network edge before reaching the application cluster.\n\nBecause the application services themselves remained up and healthy, this failure mode was not immediately visible through internal health checks - the services were isolated from receiving inbound public internet traffic rather than failing.\n\nThe root cause of why this specific switch exhibited asymmetric routing behavior following an otherwise valid configuration change remains under active investigation. Whether this was triggered by a switch platform behavior or a software version-specific bug is being determined through staging environment reproduction.\n\n## **What are we doing to prevent recurrence?**\n\n### **Immediate Actions**\n\n* **Configuration Rollback:** Upon identifying the routing anomaly, a full rollback of the network configuration was executed across all affected edge switches. Routing announcements and TCP reachability of the public production IP addresses were verified following the rollback. All affected services - DCD, Partner Portal, Reseller Portal, and IAM - were confirmed fully operational by 16:35 UTC. \\(DONE\\)\n\n### **Short-term**\n\n* **Staging Environment Replication:** Detailed test cases are being executed in the staging environment to reproduce the exact routing propagation behavior under the same switch configuration conditions. The goal is to determine whether the anomaly is attributable to a specific software version bug or a switch platform behavior, so that the precise failure condition can be isolated and addressed before any future rollout. ETA: Within two weeks\n\n### **Mid-term**\n\n* **Revised Rollout Strategy:** Based on the findings from staging, a revised rollout approach will be designed to ensure that any future application of these routing filter optimizations can be performed with greater stability guarantees - including more granular validation checkpoints between switch-level changes. ETA: October 2026\n\n## **Closing remarks**\n\nWe recognize that loss of access to IAM and DCD carries real operational impact. The fact that the underlying services were healthy throughout is not a mitigation of that impact.\n\nWe are committed to ensuring that the root cause is fully understood before any re-attempt of the original change, and that the revised rollout strategy addresses the conditions that led to the asymmetric routing behavior.\n\nWe thank you for your patience while we conclude the investigation.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-08-31T17:03:36.666Z",
"resolved_inferred": false,
"started_at": "2026-08-31T15:35:23.659Z",
"state": "postmortem",
"title": "DCD is Currently Unavailable",
"updated_at": "2026-09-01T16:52:57.570Z",
"url": "https://stspg.io/313v8jdt0dm8"
},
{
"body": "# Root Cause Analysis\n\n## What happened?\n\nStarting approximately 19:00 UTC on 24 August 2026, customers accessing S3 Object Storage in the Frankfurt \\(FRA4\\) data center experienced elevated latency across all operations - uploads, downloads, metadata requests, and deletions. Intermittent HTTP 503 Service Unavailable and 404 Not Found errors were observed on object read requests. The impact was measurable for all customer which had buckets in the both affected datacenters of the region, with some customers experiencing severe degradation depending on their bucket configuration and access patterns.\n\nThe incident remained in progress with high priority from 25 August through 3 September 2026. The latency issue was mitigated at approximately 13:40 UTC on 3 September 2026.\n\n## How was this possible? \\(Root Cause\\)\n\n**Primary cause - Software bug in Quality of Service \\(QoS\\) subsystem**\n\nIONOS S3 Object Storage in the Frankfurt region is using a distributed object storage system. A QoS feature of this service\u00a0 - implemented via a redis-qos service - applies rate limiting to S3 requests at the cluster level. A bug in the S3 service leads to not well distributed queries to the Redis-QOS service \\(which is served by multiple server for high availability\\) this caused the Redis-QOS process to reach and sustain 100% CPU utilization, progressively slowing down all S3 request processing at the cluster level. This affected every request passing through the affected nodes, irrespective of the operation type or the specific bucket accessed.\n\n\u200cThis is an internal defect within the software used. The bug caused the S3 service to consume all available resources before load reached request levels that would normally trigger throttling, meaning the degradation occurred continuously rather than only under peak conditions.\n\n\u200cIn close cooperation with the software vendor, we disabled the QoS rate-limiting function on 3 September, which fully resolved the latency issue. S3 is currently operating without some QoS features while a permanent fix is prepared by the vendor.\n\n## What are we doing to prevent recurrence?\n\n**Already completed:**\n\n* \u200cQoS disabled in the S3 service - fully resolved the latency issue. \\(DONE\\)\n* Requesting permanent solution from the software vendor. \\(INPROGRESS\\)\n\n\u200c**Short-term - ETA: within 2 weeks:**\n\n* Permanent fix for the Cloudian QoS bug: IONOS Cloud is in active coordination with the vendor to obtain and deploy a fix for the redis-qos defect. Once the fix is validated, lost QoS features will be re-enabled.\n* Database partition monitoring: We are implementing monitoring that alerts on partition size growth before any individual partition approaches a problematic threshold. This will allow our team to identify and address bucket layout issues proactively.\n\n**Mid-term - ETA: 1 to 3 months:**\n\n* QoS architecture review: Following the permanent QoS fix, we will review the architectural isolation of the QoS service together with the vendor to ensure that a future resource contention event in the rate-limiting layer cannot propagate to the request path at the same scale.\n* Monitoring and alerting improvements: We are extending cluster-level monitoring to surface redis-qos CPU saturation and database compaction backlog as first-class incident signals, with automated escalation before customer-visible latency develops.\n\n## Closing remarks\n\nAn incident of this duration in a core infrastructure service is not acceptable. The high-latency period persisted for nine days, during which customer workloads depending on S3 in the Frankfurt region were degraded. Multiple optimisations and mitigation strategies were implemented during the course of the incident, but could only improve the situation for individual buckets and only to a certain extent. Detecting the underlying QoS bug and developing a mitigation required coordination with the vendor\u2019s engineering team.\n\nWhile the latency issue is mitigated, we remain in close contact with the vendor. The engineering work to deliver a permanent QoS fix, reduce database partition pressure, and prevent recurrence is in progress. We are also working closely with our technology partner to understand delays in the analysis of the root cause of this incident. We will conduct a joint post mortem to identify areas where collaboration during incidents can be improved.\n\nWe recognise the impact this incident caused to your operations. We believe that the listed measures will help us prevent similar error patterns and speed up analysis and recovery for software related issues in the future.\n\nWe thank you for your patience during the incident.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-09-03T13:39:13.199Z",
"resolved_inferred": false,
"started_at": "2026-08-26T09:53:49.200Z",
"state": "postmortem",
"title": "Object Storage - Increased latency in eu-central-1",
"updated_at": "2026-09-09T13:58:51.304Z",
"url": "https://stspg.io/5bfcrzmykmzl"
},
{
"body": "This incident has been resolved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-08-23T17:48:12.132Z",
"resolved_inferred": false,
"started_at": "2026-08-23T12:47:45.805Z",
"state": "resolved",
"title": "Object Storage Service Restrictions",
"updated_at": "2026-08-23T17:48:12.147Z",
"url": "https://stspg.io/3z09mypzm01j"
},
{
"body": "This incident has been resolved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-08-23T17:49:25.108Z",
"resolved_inferred": false,
"started_at": "2026-08-21T07:52:17.299Z",
"state": "resolved",
"title": "IP Reservation Not Possible",
"updated_at": "2026-08-23T17:49:25.124Z",
"url": "https://stspg.io/zdfpn8nmwyfr"
},
{
"body": "We managed to get back to normal capacity, so Support is back to normal availability.\nThank you for the patience.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-08-20T08:18:47.143Z",
"resolved_inferred": false,
"started_at": "2026-08-14T14:50:50.000Z",
"state": "resolved",
"title": "Cloud Support: Telephone Support reduced capacity",
"updated_at": "2026-08-20T08:18:47.165Z",
"url": "https://stspg.io/6k509qzjs90d"
},
{
"body": "This incident has been resolved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-08-14T13:34:29.070Z",
"resolved_inferred": false,
"started_at": "2026-08-14T10:51:10.546Z",
"state": "resolved",
"title": "Provisioning: VDC-14-1836",
"updated_at": "2026-08-14T13:34:29.091Z",
"url": "https://stspg.io/w7vmxbq4yp0w"
},
{
"body": "We were able to resolve this issue, it's back to normal now.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-08-14T08:04:23.822Z",
"resolved_inferred": false,
"started_at": "2026-08-13T00:49:59.000Z",
"state": "resolved",
"title": "AI Model Hub - Service Degradations",
"updated_at": "2026-08-14T08:04:23.845Z",
"url": "https://stspg.io/07wzf22f2x01"
},
{
"body": "This incident has been resolved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-08-12T11:52:18.481Z",
"resolved_inferred": false,
"started_at": "2026-08-11T16:05:26.026Z",
"state": "resolved",
"title": "Limited access to provisioning services",
"updated_at": "2026-08-12T11:52:18.499Z",
"url": "https://stspg.io/wcfk0n37812r"
},
{
"body": "We are marking this incident as resolved.\nThe measures already put into place have had the desired effect on the service and have improved performance and stability. \nWhile this incident is marked as resolved, we are continuing executing on our action plan to improve the performance and reliability of our Managed Kubernetes Control Planes. Our current focus:\n- Further load balancing on our Control Plane Clusters\n- Migration of control planes to improved infrastructure",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-08-10T19:58:43.820Z",
"resolved_inferred": false,
"started_at": "2026-08-04T10:24:31.473Z",
"state": "resolved",
"title": "Managed Kubernetes - Intermittent Control Plane Unavailability",
"updated_at": "2026-08-10T19:58:43.844Z",
"url": "https://stspg.io/lf7frx7rltr6"
},
{
"body": "We were able to assign additional personnel and will set this status page to resolved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-08-04T10:02:32.923Z",
"resolved_inferred": false,
"started_at": "2026-08-04T08:20:14.167Z",
"state": "resolved",
"title": "Cloud Support: Telephone Line Availability Degraded",
"updated_at": "2026-08-04T10:02:32.937Z",
"url": "https://stspg.io/jzcx35pb662q"
},
{
"body": "We are marking this incident as resolved as no further abnormalities could be detected. We will share an RCA as soon as it is compiled.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-28T16:36:10.020Z",
"resolved_inferred": false,
"started_at": "2026-07-28T13:46:30.000Z",
"state": "resolved",
"title": "Connectivity Issues with DBaaS (MongoDB)",
"updated_at": "2026-07-28T16:36:10.034Z",
"url": "https://stspg.io/f975x3t3ccwm"
},
{
"body": "No more anomalies were detected.\nThe Root Cause of the loss of redundancy was identified as a memory leak in a management component. A durable mitigation was put in place to avoid recurrence. The memory leak will be addressed in an upcoming update to the management component.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-25T07:49:29.292Z",
"resolved_inferred": false,
"started_at": "2026-07-25T06:27:04.561Z",
"state": "resolved",
"title": "LAS: Storage Loss of Redundancy",
"updated_at": "2026-07-25T07:49:29.314Z",
"url": "https://stspg.io/jqd4pvg8q3k5"
},
{
"body": "Cloud Support availability is back.\nThank you for your understanding.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-27T05:43:45.267Z",
"resolved_inferred": false,
"started_at": "2026-07-24T06:51:40.324Z",
"state": "resolved",
"title": "Cloud Support: Limited Phone Support Availability",
"updated_at": "2026-07-27T05:43:45.282Z",
"url": "https://stspg.io/qwkfzw8y09ty"
},
{
"body": "We are marking this incident as resolved. Our AI Modelhub Team has addressed the issue, which was caused by acute resource constraints. This was resolved by scaling out the bottlenecks.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-28T07:12:22.698Z",
"resolved_inferred": false,
"started_at": "2026-07-23T10:44:26.165Z",
"state": "resolved",
"title": "AI Model Hub - Service Degradations",
"updated_at": "2026-07-28T07:12:22.715Z",
"url": "https://stspg.io/hv2d4dl7y6qd"
},
{
"body": "we are marking this incident as resolved as the underlying storage issue has been resolved. We will create a follow up incident for Managed Kubernetes performance and stability issues to avoid confusion.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-08-04T10:01:12.052Z",
"resolved_inferred": false,
"started_at": "2026-07-18T13:33:01.430Z",
"state": "resolved",
"title": "MK8s - Connectivity Issue",
"updated_at": "2026-08-04T10:01:12.073Z",
"url": "https://stspg.io/wq43g68kdl4w"
},
{
"body": "Phone coverage is currently normal",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-17T16:55:00.499Z",
"resolved_inferred": false,
"started_at": "2026-07-15T10:15:02.696Z",
"state": "resolved",
"title": "Cloud Support: Upcoming Limited Phone Support Availability",
"updated_at": "2026-07-17T16:55:00.522Z",
"url": "https://stspg.io/hjc15ymvgrm5"
},
{
"body": "This was a duplicate.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-13T11:21:12.796Z",
"resolved_inferred": false,
"started_at": "2026-07-13T11:18:46.387Z",
"state": "resolved",
"title": "VMs and managed services availability in FRA degraded",
"updated_at": "2026-07-13T11:21:12.810Z",
"url": "https://stspg.io/njhsjsgd5wql"
},
{
"body": "The incident has been resolved, the VMs are back working again.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-13T13:44:18.478Z",
"resolved_inferred": false,
"started_at": "2026-07-13T11:16:33.454Z",
"state": "resolved",
"title": "Resources in an Unknown state - FRA",
"updated_at": "2026-07-13T13:44:18.495Z",
"url": "https://stspg.io/p187f14ckjqz"
},
{
"body": "**What happened?**\n\nOn July 9, 2026, the Logro\u00f1o datacenter experienced a cascading power and cooling failure impacting customer workloads and services. The incident began at 21:59 CEST and required on-site manual intervention to restore service. Full power was restored at 23:05 CEST; complete service normalization concluded at 09:22 CEST on July 10.\n\nNo customer data loss has been identified.\n\n**How was this possible? \\(Root Cause\\)**\n\nThe outage was caused by a single electrical fault that triggered a cascading sequence of automated safety responses.\n\nAt 21:59 CEST, a residual current device \\(RCD breaker D1\\) in the high-voltage control cabinet tripped. At this point, no visible outage occurred. However, the trip silently disabled the generator control system - the generators lost visibility of their circuit breakers and were no longer able to start automatically. The facility appeared fully operational, but its backup power capability was gone.\n\n23 minutes later, at 22:22 CEST, the 66 kV high-voltage grid supply failed. The protection system detected overcurrent and initiated an automatic safety shutdown - a designed safety response, not a malfunction. Because the generators had been silently disabled since 21:59, they were unable to take over the load. All UPS systems switched to battery mode, protecting the workload throughout. However, as cooling infrastructure is not UPS-backed, all chillers and CRAH units lost power immediately.\n\nManual intervention by on-site staff restored grid power at 22:50 CEST. At 23:05 CEST, all generators were reset and the transformer breaker was closed, fully restoring power to the facility. The sequential service recovery process then began.\n\nThe exact cause of the initial D1 trip is still under investigation by our electrical engineering team and an external vendor performing forensic analysis on-site.\n\n**What we are doing to prevent recurrence**\n\nThe investigation has already identified the architectural root cause: a single RCD breaker acting as a shared upstream dependency for both the generator control system and the high-voltage protection system. The following measures are being implemented:\n\nAlready completed\n\n* Generator validation: All diesel generators have been confirmed synchronized and operationally ready in auto-mode following the corrective actions applied on-site.\n* Infrastructure validation: On July 13, 2026, a scheduled maintenance shutdown by grid operator Iberdrola provided a real-world validation. All systems performed as designed \u2014 generators started automatically, no service interruption occurred.\n* High-voltage station vendor inspection completed \u2014 no additional defects found.\n\nShort-term - within 4 weeks\n\n* Redundant AC supply for HV protection and generator control system \u2014 eliminating the identified single point of failure so that no single component can simultaneously disable both primary and backup power paths.\n* Review of RCD sizing and placement in the HV auxiliary circuit.\n* Network failover review: engineering is reviewing route reflector and load balancer failover behavior to prevent static connectivity conditions that compounded recovery time.\n\nStructural\n\n* Redundant AC supply for HV protection and generator control system \u2014 eliminating the identified single point of failure so that no single component can simultaneously disable both primary and backup power paths.\n* Emergency runbook update: HV protection system failure scenario added and distributed to all on-site and on-call staff.\n* Full electrical architecture audit: a comprehensive review of the Logro\u00f1o facility's electrical distribution layout to identify and remediate any further single points of failure.\n* Full post-incident review \\(PIR\\) to be completed and shared upon request.\n\nClosing remarks\n\nDespite the resilience measures already in place, service disruption could not be completely avoided during this event. The cascading nature of the failure - where a single circuit trip silently disabled backup power capability before the grid failure occurred - is a serious architectural finding, and we have moved immediately to eliminate that dependency.\n\nOn July 13, 2026, our infrastructure was put to a real-world test during a scheduled Iberdrola maintenance shutdown. All systems performed as designed. We are confident that the corrective actions taken are effective.\n\nWe recognize the impact this had on customers relying on our services and are committed to closing all identified action items with urgency and transparency.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-10T07:59:30.859Z",
"resolved_inferred": false,
"started_at": "2026-07-09T22:16:02.859Z",
"state": "postmortem",
"title": "Power Infrastructure Incident in VIT",
"updated_at": "2026-07-22T09:59:44.878Z",
"url": "https://stspg.io/tylz38855z6l"
},
{
"body": "This incident has been resolved. We are working on the RCA and will publish it here asap",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-08T09:31:39.399Z",
"resolved_inferred": false,
"started_at": "2026-07-06T10:14:04.138Z",
"state": "resolved",
"title": "Provisioning: Degraded Performance in Job Processing",
"updated_at": "2026-07-08T09:31:39.420Z",
"url": "https://stspg.io/plt79161tg5g"
},
{
"body": "This incident has been resolved. We are working on the RCA and will publish it here asap",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-08T09:33:23.393Z",
"resolved_inferred": false,
"started_at": "2026-07-05T14:47:50.000Z",
"state": "resolved",
"title": "Partial Connectivity Degradation to Control Plane Affecting Kubernetes Operations",
"updated_at": "2026-07-08T09:33:23.409Z",
"url": "https://stspg.io/w0bmkc6bdyxh"
},
{
"body": "The provisioning queue looks healthy again, and all services have been reactivated.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-05T11:50:57.640Z",
"resolved_inferred": false,
"started_at": "2026-07-05T08:01:04.215Z",
"state": "resolved",
"title": "Provisioning: Degraded Performance in Job Processing",
"updated_at": "2026-07-05T11:50:57.672Z",
"url": "https://stspg.io/f4ft6txlk3d4"
},
{
"body": "Our Provisioning went back operational yesterday around 16:30 UTC.\nFurther monitoring showed no issues anymore, so this incident was resolved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-03T09:16:05.179Z",
"resolved_inferred": false,
"started_at": "2026-07-02T08:30:41.103Z",
"state": "resolved",
"title": "Provisioning issues in FRA and TXL",
"updated_at": "2026-07-03T09:16:05.198Z",
"url": "https://stspg.io/6c5df54jsdl1"
},
{
"body": "Phone Support is now available again",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-02T13:50:46.674Z",
"resolved_inferred": false,
"started_at": "2026-07-02T07:00:33.643Z",
"state": "resolved",
"title": "Cloud Support: Telephone Support temporarily unavailable",
"updated_at": "2026-07-02T13:50:46.690Z",
"url": "https://stspg.io/fnp7d253svq7"
},
{
"body": "This incident has been resolved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-07T10:09:05.049Z",
"resolved_inferred": false,
"started_at": "2026-06-30T21:08:11.735Z",
"state": "resolved",
"title": "Partial Control Plane Outage Affecting Kubernetes Operations",
"updated_at": "2026-07-07T10:09:05.067Z",
"url": "https://stspg.io/qlwl7yryfxgk"
},
{
"body": "This incident has been resolved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "critical",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-06-23T14:26:14.805Z",
"resolved_inferred": false,
"started_at": "2026-06-23T10:24:00.000Z",
"state": "resolved",
"title": "AI Model Hub - All models currently unavailable",
"updated_at": "2026-06-23T14:26:14.820Z",
"url": "https://stspg.io/d77qt9592dws"
},
{
"body": "We have enabled maintenance again in all DCs.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-08-25T17:24:26.576Z",
"resolved_inferred": false,
"started_at": "2026-06-22T19:07:02.887Z",
"state": "resolved",
"title": "Managed Kubernetes - Automated Maintenance disabled",
"updated_at": "2026-08-25T17:24:26.592Z",
"url": "https://stspg.io/yhx9n9d9lqrf"
},
{
"body": "We have marked this incident as resolved. Several underlying VMs entered an unhealthy state, which increased the CPU load on the remaining healthy VMs supporting the service. This cascaded into the intermittent availability issues. All systems are now operating normally.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-06-22T11:04:25.210Z",
"resolved_inferred": false,
"started_at": "2026-06-22T08:01:54.696Z",
"state": "resolved",
"title": "MK8s Control Plane: Partial Outage",
"updated_at": "2026-06-22T11:04:25.227Z",
"url": "https://stspg.io/f2s7zr9yjtb9"
},
{
"body": "We are marking this incident as resolved, as no further abnormalities occurred.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-06-22T18:42:14.514Z",
"resolved_inferred": false,
"started_at": "2026-06-19T06:35:14.754Z",
"state": "resolved",
"title": "S3: Increased Error Rate",
"updated_at": "2026-06-22T18:42:14.529Z",
"url": "https://stspg.io/g56xxff576vl"
},
{
"body": "We are resolving this incident as provisioning performance has fully normalized.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-06-19T16:32:33.183Z",
"resolved_inferred": false,
"started_at": "2026-06-18T17:20:03.000Z",
"state": "resolved",
"title": "Provisioning: Degraded Performance in Job Processing",
"updated_at": "2026-06-19T16:32:33.198Z",
"url": "https://stspg.io/hl59n267w7xs"
},
{
"body": "We are currently experiencing exceptionally high demand for GPU servers, which has temporarily limited our available capacity. As a result, you may encounter provisioning errors when attempting to create a new GPU server or start an existing one.\n\nWhat You Can Do\n\n- Try again later: We cannot provide an exact timeline for when capacity will be added. However, we are monitoring the situation closely and will update this status page once substantial free capacity becomes available.\n- Provision a smaller GPU server: Because availability fluctuates rapidly, we cannot guarantee which specific sizes are currently open. Attempting to provision a smaller instance may increase your chances of success.\n- If you currently have a provisioned and running GPU server, do not shut it down. Due to the low supply, freed resources may be immediately allocated to other customers, and you will likely be unable to restart your server.\n\nNext Steps & Resolution\nResolving this resource constraint has high priority. Our teams are working to increase capacity to meet demand as quickly as possible.\n\nWe sincerely appreciate your patience and understanding.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"started_at": "2026-06-18T15:04:07.611Z",
"state": "identified",
"title": "GPU Server Provisioning: Supply currently limited",
"updated_at": "2026-06-18T15:04:07.885Z",
"url": "https://stspg.io/7nppvj7b5bbr"
},
{
"body": "The incident has now been resolved, and accessing S3 Buckets should now be working as expected.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-06-17T14:03:07.052Z",
"resolved_inferred": false,
"started_at": "2026-06-17T11:40:18.532Z",
"state": "resolved",
"title": "S3 High Error Rate - TXL",
"updated_at": "2026-06-17T14:03:07.067Z",
"url": "https://stspg.io/59ns6vzf7qgn"
},
{
"body": "The system is stable. We are resolving this incident. The trigger of this incident will be shared once the analysis has been completed.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-06-15T11:20:45.560Z",
"resolved_inferred": false,
"started_at": "2026-06-15T09:03:31.209Z",
"state": "resolved",
"title": "Availability Issue DCD",
"updated_at": "2026-06-15T11:20:45.595Z",
"url": "https://stspg.io/sn5lcg9d5vfk"
},
{
"body": "Phone Support is now available again",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-06-13T21:13:38.945Z",
"resolved_inferred": false,
"started_at": "2026-06-13T13:50:27.797Z",
"state": "resolved",
"title": "Cloud Support: Telephone Support temporarily unavailable",
"updated_at": "2026-06-13T21:13:38.963Z",
"url": "https://stspg.io/2nj892jxvgy0"
},
{
"body": "MongoDB Playground and Business Edition clusters are currently not available for provisioning in de/fra/2 (Frankfurt East). This is due to a capacity limitation in that location that prevents these cluster types from being created reliably, and it is expected to persist for the time being. We will update this status page as soon as this limitation is lifted.\n\nAs an alternative, you can:\n- Provision your Playground or Business Edition cluster in another available location, or\n- Use MongoDB Enterprise Edition, which remains available in Frankfurt East.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"started_at": "2026-06-12T15:46:42.244Z",
"state": "investigating",
"title": "Provisioning of MongoDB Playground and Business Edition Clusters unavailable",
"updated_at": "2026-06-12T15:46:42.356Z",
"url": "https://stspg.io/745hmcys28ty"
},
{
"body": "We are closing the incident as no further error spikes were recorded. We will share the trigger of this incident once the analysis has been conducted.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-06-12T13:49:21.393Z",
"resolved_inferred": false,
"started_at": "2026-06-12T07:44:39.843Z",
"state": "resolved",
"title": "S3: eu-central-3 Increased Error Rate",
"updated_at": "2026-06-12T13:49:21.408Z",
"url": "https://stspg.io/ktgn7k54rznb"
},
{
"body": "The Object Storage endpoints are stable now.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-06-08T14:33:54.712Z",
"resolved_inferred": false,
"started_at": "2026-06-08T13:21:39.414Z",
"state": "resolved",
"title": "Object Storage unavailable in Frankfurt",
"updated_at": "2026-06-08T14:33:54.729Z",
"url": "https://stspg.io/vhs61j7hkqnk"
},
{
"body": "We are marking this incident as resolved. To address the described performance and stability concerns, we have made the following changes:\n- Upgraded our AI Modelhub setup to increase overall model performance.\n- Rolled out improved monitoring and response tooling for high-load scenarios.\n- Replaced Llama 405B with Qwen 3.5 397B.\n- Officially deprecated Collections.\n\nWe are confident these changes address these issues effectively and will help us provide an even better service moving forward. \n\nThank you for your patience!",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "critical",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-28T12:19:25.450Z",
"resolved_inferred": false,
"started_at": "2026-06-04T14:34:16.000Z",
"state": "resolved",
"title": "AI Model Hub - Service Degradations",
"updated_at": "2026-07-28T12:19:25.466Z",
"url": "https://stspg.io/h02nkwhj8212"
},
{
"body": "This incident has been resolved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-06-02T11:28:51.073Z",
"resolved_inferred": false,
"started_at": "2026-06-01T14:10:08.285Z",
"state": "resolved",
"title": "Limited access to provisioning services",
"updated_at": "2026-06-02T11:28:51.093Z",
"url": "https://stspg.io/rvmxd10w0llp"
},
{
"body": "This incident has been resolved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-06-01T17:06:36.465Z",
"resolved_inferred": false,
"started_at": "2026-05-29T20:23:33.410Z",
"state": "resolved",
"title": "AI Model Hub: Increased Error Rate",
"updated_at": "2026-06-01T17:06:36.483Z",
"url": "https://stspg.io/kfvsmqyl4l2y"
},
{
"body": "We are marking this incident as resolved. Our team is currently investigating a link to a previous maintenance task on the VPN gateways. We will update the incident with the investigation results once they are available.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-05-28T17:21:58.797Z",
"resolved_inferred": false,
"started_at": "2026-05-28T10:43:21.086Z",
"state": "resolved",
"title": "Connectivity Degradation: VPN Gateways",
"updated_at": "2026-05-28T17:21:58.813Z",
"url": "https://stspg.io/c54yllkcsh5x"
},
{
"body": "In this RCA we want to share our analysis of a network connectivity incident that affected virtual servers on 20 May 2026 in our Worcester datacenter. During this period, LANs associated with a subset of your virtual machines became unavailable due to a network configuration issue in the underlying cluster. We have documented the sequence of events, the technical root cause, and the measures we are implementing to prevent recurrence.\n\n**What happened?**\n\nOn 20 May 2026 at approximately 11:19 UTC, LAN connectivity was lost for virtual machines hosted on a subset of physical servers in an affected cluster. The impact manifested as a complete loss of LAN connectivity for the virtual machines running on those servers.\n\nEngineering was alerted and began investigation. The root cause was identified, and the affected virtual machines were migrated to alternative physical servers with compatible interface speeds. LAN connectivity was fully restored by 14:27 UTC.\n\n**How was that possible? \\(Root Cause\\)**\n\nThe affected cluster contained a mix of physical servers operating at different interface speeds: a subset of servers were running at 100Gbps, while the remainder of the cluster had been upgraded to 200Gbps. This mixed-speed configuration was present prior to the incident.\n\nMulticast groups - which underpin LAN connectivity between virtual machines - form at the interface speed of the highest-speed server participating at the time the group is created. By platform design, a server may join an existing group only if its interface speed is equal to or greater than the group's speed. A faster server can therefore join a slower group, and its presence does not raise the group's speed, but a slower server cannot join a faster group.\n\nBefore the incident, the affected groups had been created at 100Gbps. The 200Gbps servers were able to participate in them because their speed exceeded the group speed, and their presence did not change it. As a result, LANs spanning both 100Gbps and 200Gbps hosts operated normally.\n\nOn 20 May, a platform-level liveboot rollout of a gateway component caused multicast groups to be recreated across the cluster. During recreation, each group's speed was re-derived from the highest-speed server then participating. Any group spanning both 100Gbps and 200Gbps hosts therefore reformed at 200Gbps. The 100Gbps servers were now slower than the group and were rejected on rejoin. All multicast join requests from the affected 100Gbps servers failed, and the LANs hosted on those servers went down.\n\nThe gateway liveboot was the operational trigger that caused multicast groups to be rebuilt. The underlying condition that made those rebuilds produce a connectivity failure was the presence of servers with 100Gbps interface speeds in an otherwise 200Gbps cluster. This disparity had not been detected prior to the event.\n\n**What we are doing to prevent recurrence**\n\nCompleted\n\n* The affected virtual machines were migrated to physical servers operating at 200Gbps, restoring LAN connectivity. The 100Gbps servers have been removed from active customer workloads.\n\nIn progress\n\n* Interface upgrades: The affected physical servers are being upgraded to 200Gbps network interface cards, bringing their interface speed in line with the rest of the cluster. This eliminates the mixed-speed condition entirely. \\(ETA: Q3 2026\\)\n* Cluster speed monitoring: We are implementing automated monitoring to detect physical servers operating at non-standard interface speeds within a cluster. This will ensure that any future speed disparity is identified and flagged before it can affect customer workloads during operational events such as gateway liveboots. \\(ETA: Q3 2026\\)\n\nStructural review\n\nWe are reviewing our gateway liveboot and rollout procedures to assess whether multicast group recreation can be performed in a way that is resilient to mixed-speed cluster configurations, adding a safeguard independent of monitoring coverage.\n\n**Closing remarks**\n\nThe loss of LAN connectivity, even for a bounded period, disrupts operational continuity, and we recognize the impact this incident had on your environment. The root condition - a speed disparity between servers in the same cluster - should have been identified before it could be exposed by a routine operational event. We have addressed the immediate impact, and the NIC upgrades and monitoring improvements underway will close the gap that allowed this configuration to persist undetected.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-05-20T21:38:59.374Z",
"resolved_inferred": false,
"started_at": "2026-05-20T12:19:10.657Z",
"state": "postmortem",
"title": "Network Outage in Worcester",
"updated_at": "2026-06-15T19:22:08.612Z",
"url": "https://stspg.io/b2z1g2wt3d33"
},
{
"body": "We are marking this incident as resolved. \nThe root cause of the service degradation affecting Kubernetes Control Planes in FRA has been identified as a DDoS attack targeting customers hosted on our infrastructure. This incident has been documented here:\nhttps://status.ionos.cloud/incidents/kknrz6284jl0\n\nThe identified bottleneck led to the observed intermittent reachability issues and increased latency between Kubernetes etcd nodes for the affected customers.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-05-19T16:24:31.143Z",
"resolved_inferred": false,
"started_at": "2026-05-15T12:18:48.543Z",
"state": "resolved",
"title": "Partial Connectivity Degradation to Control Plane Affecting Kubernetes Operations in Frankfurt (de/fra)",
"updated_at": "2026-05-19T16:24:31.161Z",
"url": "https://stspg.io/mxyflpmvkxlk"
},
{
"body": "We are marking this incident as resolved. \nThe Root Cause of the incident was a DDOS attack on customers hosted in our infrastructure. While the attack was largely successfully mitigated by our countermeasures a intermittent saturation of an affected link resulted in an increase of package loss.\n\nOur network team has evaluated the countermeasures' effectiveness and the identified bottleneck and is currently working on technical and organizational improvements to further decrease detection time for comparable events.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-05-19T16:17:18.711Z",
"resolved_inferred": false,
"started_at": "2026-05-15T11:59:31.105Z",
"state": "resolved",
"title": "Network packet loss",
"updated_at": "2026-05-19T16:17:18.727Z",
"url": "https://stspg.io/krd8z5c638r2"
},
{
"body": "Redundancy has been fully restored and no limitations related to provisioning should remain.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-05-18T07:36:29.638Z",
"resolved_inferred": false,
"started_at": "2026-05-13T15:33:46.516Z",
"state": "resolved",
"title": "Provisioning Service Degradation in FKB",
"updated_at": "2026-05-18T07:36:29.654Z",
"url": "https://stspg.io/wlzq8kdb2035"
},
{
"body": "This incident has been resolved, and all affected customers have been contacted directly via email.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-08T12:29:58Z",
"resolved_at": "2026-05-13T06:56:17.673Z",
"resolved_inferred": false,
"started_at": "2026-05-12T08:50:42.075Z",
"state": "resolved",
"title": "Suspected Issues with USt- / VAT ID verification",
"updated_at": "2026-05-13T06:56:17.690Z",
"url": "https://stspg.io/ss96py5844bn"
},
{
"body": "**Root Cause Analysis**\n\nBetween 11 and 12 May 2026, a cluster in our FKB datacenter was affected by two consecutive storage hardware failures, which led to service degradation and I/O outage for affected workloads spanning over 15 hours and limiting provisioning functionality in affected VDCs.\n\nIn the following Root Cause Analysis, we explain the incident, identify the technical triggers of the outage, describe the work done during mitigation, and outline the measures we are taking to prevent a recurrence.\n\n**Technical Context**\n\nBecause the failure mode was a cascading sequence of two hardware events on a redundant storage pair, we want to provide the technical context necessary to understand how a redundant system could be affected to this degree:\n\nTo provide a resilient service, storage servers in the affected clusters are deployed in redundant pairs \\(Leg A and Leg B\\) that hold synchronized data. Each leg itself has built-in redundancy and fault tolerance  - for example, via RAID configurations for disks.\n\nThe two legs are designed to mitigate risk from hardware failures or issues that would render the resilience mechanisms of a single leg ineffective. Failure of one leg is largely transparent to the VM host, which - for customers - continues to operate normally.\n\nIn a \"Zero Leg\" event - where connectivity to both legs is lost - VMs served by the affected storage experience an outage. This is the event that caused this incident.\n\n**What Happened**\n\nOn 11 May 2026 , the first of two servers \\(Leg A\\) in a redundant storage backend pair began reporting multiple disk errors in its RAID array. The error pattern - UDMA CRC errors across several disks - was consistent with degradation of a shared physical interconnect \\(SAS backplane or cabling\\), rather than wear on any single disk. From this point onward, all dependent storage volumes were operating without their secondary failover path, served exclusively from Leg B - a single-leg event.\n\nAs a forced array reassembly was initiated  additional disks faulted during the rebuild in the afternoon, degrading the array below the recoverable threshold. A subsequent reboot brought the array up fully inactive, and Leg A was considered unrecoverable. The decision was made to begin migrating redundancy onto alternative storage systems, with the goal of restoring two-leg operation. Due to the volume of data that needed to be synchronized, this process was expected to take several hours.\n\nOn 12 May at 05:13 UTC - while synchronization was still in progress - Leg B experienced an independent disk disconnection in its own RAID array. With Leg A already non-functional, this second failure eliminated all remaining redundancy in the pair. All dependent volumes dropped to zero available copies simultaneously, causing an I/O outage for the affected virtual machines.\n\nLeg B was rebooted at 05:38 UTC and its array force-assembled with the remaining healthy disks, which allowed a first wave of volume recovery to complete by 08:25 UTC. Additional disk failures on the same server then materialized at 10:59 UTC, triggering a second outage. At this point, the decision was taken to perform a complete server replacement: at approximately 14:30 UTC, all disks were physically transferred to new server hardware and the RAID arrays were rebuilt on the replacement system.\n\nThe hardware swap introduced an additional recovery complication: the replacement server now carried a different storage network identity. Compute nodes across the cluster still held cached connection sessions pointing to the original hardware identity, and VM restart attempts subsequently failed. When a cleanup of stale connection mappings across affected compute nodes in the fleet was performed existing automation tooling proved ineffective. This significantly hampered recovery efforts. Leg B was confirmed fully operational at 20:45 UTC, and affected VMs began recovering. VMs where automatic recovery was not possible were remediated manually in the following hours.\n\nDuring the subsequent synchronization process to re-establish redundancy between the new legs, the provisioning service was deactivated in VDCs that contained VMs served by the affected storage pair. This meant customers were temporarily unable to roll out changes to affected VMs. Storage redundancy was fully restored on 15 May, once all affected volumes had been migrated back to a fully redundant configuration. At this point, the service was fully recovered.\n\n**How was that possible?\\(Root Cause\\)**\n\nThe incident was caused by two independent hardware failures on the two servers of a single redundant storage pair, occurring several hours apart. Each failure on its own would not have caused customer impact. The combination of both, within the window between the first failure and the completion of redundancy migration, eliminated the redundancy that the architecture is designed to provide.\n\n_Primary failure - Leg A._ The errors observed on Leg A were consistent with degradation of a shared physical interconnect - a SAS backplane or cabling issue - rather than end-of-life wear on individual disks. This pattern explains both why multiple disks failed in a correlated way, and why repair attempts could not outpace degradation: the rebuild operations themselves exercise the interconnect, and a second disk faulted under rebuild load before the array could complete. A stale array member prevented full reassembly, and the server was effectively out of service after its reboot.\n\n_Secondary failure - Leg B._ While migration of redundancy to alternative storage was in progress, Leg B suffered an independent disk disconnection. The exact root cause of this second failure is still under investigation, but the symptoms were consistent with previously observed wear-related failures within the same hardware generation and make.\n\n_Recovery complication - network identity mismatch._ When the disks from Leg B were transferred into a replacement chassis, the new server presented a different storage network identity. Compute nodes in the cluster, however, retained cached connection sessions pointing to the original hardware identity. VM restart attempts therefore failed against the replacement hardware until those sessions were invalidated across every affected compute node. As available automation tooling proved ineffective specialized tooling needed to be developed and tested during the incident.\n\nIn summary: a single-leg operation window - opened by the first hardware failure - coincided with an independent second hardware failure on the surviving leg before redundancy could be restored. The recovery path then required full hardware replacement, which surfaced inadequate automation in the storage-to-compute identity handover which prolonged the recovery process significantly.\n\n**What we are doing to prevent recurrence**\n\nWe have grouped the remediation into completed tasks, immediate measures, and a mid-term reviews to address the gaps that turned a localized hardware failure into a prolonged customer-impacting outage.\n\n_Already taken:_\n\n* **Complete server replacement:** All disks from the affected servers were physically transferred to verified replacement hardware, with RAID arrays fully rebuilt and verified on the new platform.\n* **Disk replacement and array verification:** All involved disks were replaced; arrays were rebuilt and integrity-verified before being returned to service.\n\n_Immediate \\(June 2026\\):_\n\n* **Automated session migration on hardware replacement:** Update and validate tooling that automatically refreshes storage network identity mappings across all compute nodes when a storage server of this configuration is replaced in the manner required during this incident. This directly addresses the cleanup step that extended recovery by several hours.\n* **Single-leg exposure review:** Tighten the operational window during which a degraded redundant pair is allowed to operate on a single leg in configurations in locations where Dynamic Leg Swapping is not available. Dynamic Leg Swapping refers to a progress where a single leg failure initiates an automated sync to alternate storage. Repair on the affected leg can continue and depending on which redundancy recovery strategy is quicker \\(repair vs migration\\) the leg that is available quickest will be chosen. This bounds the time spent on repair efforts, simplifies decision-making during incident response, and shrinks the window in which a subsequent hardware failure can directly impact connected workloads.\n\n_Mid-term \\(within Q3\\):_\n\n* **Dynamic Leg Swapping Rollout:** We assess the technical prerequisites to retrofit locations to support Dynamic Leg Swapping or provide alternative strategies to shorten \\(automated\\) redundancy restoration in data centers.\n* **Recovery Tool Assessment**:: The incident has shown that existing recovery tooling needs to be reviewed and validated to work against all existing hardware and configuration combinations in the fleet.\n* **Hardware generation review:** Comprehensive review of the affected storage hardware generation, covering firmware versions, backplane integrity, and environmental conditions. Failure patterns will be evaluated to rule out problematic combinations. Components showing early signs of degradation will be proactively replaced or upgraded.\n\n**Closing Remarks**\n\nWe recognize that this incident produced a full I/O outage for the virtual machines in environments that depended on this storage pair, lasting over 15 hours, and that volumes operated without their secondary failover path for several hours before and after the incident. For any production workload, an outage and loss of redundancy of this duration is significant and unacceptable.\n\nThis incident not only resulted in service disruption and subsequent degradation, but also led to significant effort for partners and customers working on their side to mitigate the resulting outage.\n\nWe believe that the measures we have completed and have planned will effectively reduce the risk of a similar failure pattern. The planned review will help to identify and reduce risk factors further. The assessment of recovery tooling will help us recover quicker and increase the robustness of our incident response.\n\nThank you for your patience during the outage and recovery and your engineering teams' constructive coordination and support throughout this event.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-05-18T07:38:46.862Z",
"resolved_inferred": false,
"started_at": "2026-05-12T06:15:58.275Z",
"state": "postmortem",
"title": "Availability of storage service partially limited in Karlsruhe",
"updated_at": "2026-05-21T11:57:59.917Z",
"url": "https://stspg.io/1mwkwq60wt0p"
},
{
"body": "This incident has been resolved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-08T12:29:58Z",
"resolved_at": "2026-05-11T21:16:29.650Z",
"resolved_inferred": false,
"started_at": "2026-05-11T17:53:31.601Z",
"state": "resolved",
"title": "Partial Connectivity Degradation to Control Plane Affecting Kubernetes Operations",
"updated_at": "2026-05-11T21:16:29.668Z",
"url": "https://stspg.io/j0kqggx63144"
},
{
"body": "**What happened?**\n\nOn May 8, 2026 between 15:22 and 16:12 \\(UTC\\+2\\) a routine internal operation to onboard a new Managed Kubernetes development cluster caused the deployment automation system to unintentionally target and modify additional clusters and install applications meant for the development cluster.\u00a0\n\nThe applications deployed to affected clusters could not be executed, as the required image pull secrets were not present in those environments, preventing any images from being pulled or run. Once the full scope of the incident was established, a recovery script was developed, validated against our staging environment, and executed in production. All unintended changes have been identified and removed by May 15, 2026.\n\n**How was this possible? \\(Root Cause\\)**\n\nThe incident was caused by a software bug in our platform framework onboarding operator, compounded by the absence of a configuration completeness gate in the deployment pipeline.\n\nThe onboarding operator is responsible for registering clusters with our deployment system. It relies on a scoping configuration file to restrict its operations to a defined set of target clusters. As part of the planned onboarding of the new development cluster, this configuration file was to be automatically propagated to the deployment branch. The propagation failed silently due to a merge conflict - the file never landed on the target branch.\n\nSimultaneously, a second change removed the onboarding operator from the exclusion list, triggering an immediate deployment. As the deployment system was not configured properly to react to missing configurations - rather than halting the sync - the operator was deployed without any scoping restriction. In this unconstrained state, it treated further cluster-api based clusters as valid onboarding targets.\n\nThe root cause is a structural gap in the CI/CD pipeline: the mechanisms to verify that all required configuration files had been successfully delivered to the target branch before a deployment was permitted proved inadequate. The cherry-pick failure was logged but did not block the pipeline. This made it possible to deploy a fully functional but unscoped operator under conditions that appeared normal to the team monitoring the rollout.\n\n**What are we doing to prevent recurrence?**\n\nUpon detection, the two immediate technical failures were resolved. The onboarding operator has been corrected and the deployment system configuration has been hardened. In parallel, the recovery effort - which included a full staging test cycle before any production changes - identified and cleaned up all affected clusters within the week following the incident.\n\n**Immediate Technical Actions**\n\n* Operator Scope Restriction: The onboarding operator logic has been corrected to require explicit namespace configuration before it can act on any clusters. It can no longer run unscoped against cluster-api based clusters by default. \\(DONE\\)\n* Deployment Configuration Safety: The deployment system has been reconfigured to fail deployments when required configuration files are absent, replacing the previous behavior that allowed deployments to proceed silently with an incomplete configuration. \\(DONE\\)\n* Cluster Cleanup: A recovery script was developed, validated in staging, and executed across all affected production clusters. All unintended installations and updates have been removed from all affected clusters. \\(DONE, completed May 15, 2026\\)\n\n**Structural Improvements:**\n\n* Pipeline Configuration Gate: We are implementing a mandatory verification step in the CI/CD pipeline to confirm that all required configuration files have been successfully propagated to the target branch before any deployment is triggered. \\(ETA: Q3 2026\\)\n* Change Management Controls: We are enforcing additional mandatory peer review and change management approvals for all deployments that touch production systems independent of change size, closing the process gap that allowed this change to proceed without the required _additional_ oversight. This is done in addition to the already established Change Approval Board review process. \\(ETA: Q3 2026\\)\n\n**Procedural Improvements**\n\nAffected customers were not properly kept informed about the cleanup efforts which compounded the uncertainty. Improvements to the incident communication process are currently being implemented into the wider incident management process ensuring that scalable options exist for tech teams to inform individually affected customers in a timely manner about the status of post-incident remediation efforts. \\(ETA: Q3 2026\\)\n\n**Closing remarks**\n\nWe recognize that unintended modification of cluster configurations represents a failure in the controls that should prevent our internal operations from influencing customer environments.\n\nThe fixes now in place and planned will close the immediate technical and procedural shortcomings. The structural measures will further harden the Change Process reducing the likelihood and potential fallout of maintenance related incidents further. \n\nThank you for your patience during the incident and the remediation efforts.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-04T15:18:55Z",
"resolved_at": "2026-05-11T08:30:41.395Z",
"resolved_inferred": false,
"started_at": "2026-05-08T15:15:06.825Z",
"state": "postmortem",
"title": "K8s Control Planes not available",
"updated_at": "2026-06-15T09:32:55.475Z",
"url": "https://stspg.io/y8g8b535pwz5"
}
]
}