{
"vendor": "Cycode",
"slug": "cycode",
"platform": "statuspage",
"status_url": "https://status.cycode.com",
"last_checked": "2026-09-16T12:28:20Z",
"last_state": "ok",
"history_backfilled": true,
"first_watched": "2026-09-04T07:06:16Z",
"incidents": [
{
"body": "Resolved ",
"first_seen": "2026-09-13T12:30:53Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-09-13T18:53:46Z",
"resolved_inferred": false,
"started_at": "2026-09-13T10:36:50Z",
"state": "resolved",
"title": "GitHub incident can affect Cycode",
"updated_at": "2026-09-13T18:53:46Z"
},
{
"body": "MemoryDB metrics are stable, the processing lag has recovered, and the system is now fully operational.",
"first_seen": "2026-09-09T12:30:34Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-09-08T17:26:17Z",
"resolved_inferred": false,
"started_at": "2026-09-08T16:53:36Z",
"state": "resolved",
"title": "Application and scanning degraded performance",
"updated_at": "2026-09-08T17:26:17Z"
},
{
"body": "The system is back to being fully operational.",
"first_seen": "2026-09-08T12:29:58Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-09-08T10:52:14Z",
"resolved_inferred": false,
"started_at": "2026-09-08T09:53:50Z",
"state": "resolved",
"title": "Application and scanning degraded performance",
"updated_at": "2026-09-08T10:52:14Z"
},
{
"body": "GitHub component: Pull Requests\nOriginal GitHub incident: https://stspg.io/0xbhhq84v4mt",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-09-01T16:05:06Z",
"resolved_inferred": false,
"started_at": "2026-09-01T15:10:38Z",
"state": "resolved",
"title": "GitHub incident can affect Cycode",
"updated_at": "2026-09-01T16:05:06Z"
},
{
"body": "GitHub component: Pull Requests\nOriginal GitHub incident: https://stspg.io/nrbwjftcz72d",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-08-27T00:30:26Z",
"resolved_inferred": false,
"started_at": "2026-08-26T23:04:59Z",
"state": "resolved",
"title": "GitHub incident can affect Cycode",
"updated_at": "2026-08-27T00:30:26Z"
},
{
"body": "The system is back to being fully operational.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-08-20T16:36:59Z",
"resolved_inferred": false,
"started_at": "2026-08-20T14:52:37Z",
"state": "resolved",
"title": "Scan Performance Degredation",
"updated_at": "2026-08-24T12:53:36Z"
},
{
"body": "**1. Summary**\n\nOn August 19, 2026, some customers experienced slower page loading and data display, as well as delays when starting or completing pull request scans. \n\nThe disruption was caused by a managed memory database entering a repeated restart cycle and not recovering automatically. This created a temporary slowdown in processing. Expected service behavior was restored after corrective actions, including replacement and increased capacity for the affected service. The provider is continuing to investigate why the service was unable to return to normal operation automatically.\n\n**2. Key Timeline (IDT)**\n\n\u2022 August 19, 2026, 16:10 IDT: We identified performance degradation affecting parts of the platform and pull request scanning.\n\n\u2022 August 19, 2026, 16:30 IDT: An initial corrective change was applied and service behavior was monitored.\n\n\u2022 August 19, 2026, 16:46 IDT: We confirmed that some delays were continuing and expanded the investigation.\n\n\u2022 August 19, 2026, 17:14 IDT: We began working with our service provider to investigate instability in the managed memory database.\n\n\u2022 August 19, 2026, 18:46 IDT: Recovery actions were in progress, with temporary mitigations in place to reduce customer impact.\n\n\u2022 August 19, 2026, 19:06 IDT: The managed memory database capacity update completed, and platform responsiveness and pull request scanning returned to expected behavior.\n\n**3. Root Cause**\n\nThe managed memory database entered a repeated restart cycle and was unable to recover automatically. This caused delays in the processing systems that support platform responsiveness and pull request scans.\n\nThe affected service underwent a routine replacement process, followed by a capacity increase during the recovery effort. While the capacity increase completed successfully and restored expected behavior, AWS is still researching why the service did not come back up normally after the restart cycle. Their detailed root-cause analysis is pending.\n\n**4. Actions Taken**\n\n\u2022 Applied corrective updates to reduce immediate processing demand.\n\n\u2022 Temporarily adjusted the status update process to reduce reliance on the affected service.\n\n\u2022 Restarted affected processing components to restore responsiveness.\n\n\u2022 Worked with the service provider to replace affected capacity and complete a capacity increase.\n\n\u2022 Monitored platform performance and scan processing until expected behavior was confirmed.\n\n**4b. Action Items**\n\n\u2022 Review the service provider\u2019s detailed root-cause analysis when available.\n\n\u2022 Improve monitoring to detect similar recovery delays earlier.\n\n\u2022 Add safeguards to reduce the impact of temporary processing slowdowns.\n\n\u2022 Introduce separate Redis capacity for each service to isolate workloads and reduce the impact of an issue in one service on others.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-08-19T18:06:00Z",
"resolved_inferred": false,
"started_at": "2026-08-19T13:18:38Z",
"state": "resolved",
"title": "Platform and PR scans slowness",
"updated_at": "2026-08-20T19:04:22Z"
},
{
"body": " \n\nSome GitHub pull request status checks in the EU region were delayed by up to 30 minutes before completing. \n\nNo checks were lost and no action is needed - delayed checks completed automatically. We identified the cause, and fixed it. We're monitoring closely.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-08-18T08:31:00Z",
"resolved_inferred": false,
"started_at": "2026-08-18T14:34:21Z",
"state": "resolved",
"title": " Delayed GitHub PR status checks",
"updated_at": "2026-08-18T14:34:21Z"
},
{
"body": "GitHub component: Git Operations, Webhooks, API Requests, Pull Requests\nOriginal GitHub incident: https://stspg.io/y1fl26l6wpzr",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-08-17T21:21:14Z",
"resolved_inferred": false,
"started_at": "2026-08-17T13:44:06Z",
"state": "resolved",
"title": "GitHub incident can affect Cycode",
"updated_at": "2026-08-17T21:21:14Z"
},
{
"body": "**Summary**\n\nPR secret scans experienced processing lag, causing delays in processing pull requests.\n\nThe issue was caused by timeouts when calling the model validation service. The model validation service was experiencing elevated latency and errors because compute nodes were scheduled on a subnet that had exhausted its available IPs.\n\nThis slowed PR secret scan processing and caused consumer lag to build up.\n\nTo resolve the issue, we temporarily disabled the feature that produces new PR secret scans so the queue could drain. We also resolved the underlying infrastructure issue by moving nodes to subnets with available free IPs.\n\n**Customer Impact**\n\nCustomers experienced delayed PR secret scans.\n\nOther scan types remained operational.\n\n**Root Cause**\n\nThe secret scanner calls the AI model validation service to classify detections during PR scans.\n\nThe model validation call had a 30-second timeout configured. When the model validation service became slow or unresponsive, each failed request could block a scan consumer for the full 30 seconds before timing out.\n\nThe underlying issue was in the model-serving infrastructure. One of the private subnets had run out of available IPs, and new nodes were created in that exhausted subnet, leaving model-serving pods in a pending state.\n\nThis resulted in elevated latency and errors in the model validation service and reduced the processing capacity of PR secret scans.\n\n**Contributing Factors**\n\n\u2022 **Model validation timeouts.** The model validation service timeout was 30 seconds. When the service was degraded, scan consumers waited for the full timeout before the request failed, which created a processing bottleneck.\n\n\u2022 **No fallback to bypass AI validation for PR scans.** There was no mechanism to skip AI model validation specifically for PR secret scans while allowing the rest of the scan to continue. The available kill switch stopped producing PR secret scans entirely. This meant the mitigation required a tradeoff between allowing processing lag to continue or temporarily stopping PR secret scans.\n\n\u2022 **Customer communication gap.** Internal alerts for PR scan lag did fire, but the status page was not updated promptly to reflect the degradation. As a result, there was a period in which we had internal visibility into the issue while customers did not have the same visibility through the status page.\n\n**Mitigation and Recovery**\n\nA pre-existing feature flag was used to stop new PR secret scans from being produced.\n\nThis stopped new messages from being added to the PR secret scan queue and allowed the existing backlog to drain.\n\nThe backlog was cleared shortly after the mitigation was enabled.\n\nIn parallel, the underlying model-serving infrastructure issue was resolved by moving nodes to subnets with available free IPs, restoring model-serving capacity.\n\nOnce the infrastructure issue was resolved and the backlog had drained, PR secret scanning was restored.\n\n**Learnings and Corrective Actions**\n\n\u2022 **Review the AI model validation timeout.** The 30-second model validation timeout meant that when the model-serving layer degraded, each scan consumer could remain blocked for the full timeout before failing. We are reviewing the timeout and retry behavior. This work must also account for the risk that shorter timeouts and faster retries could increase request volume against an already degraded model-serving service.\n\n\u2022 **Add a graceful degradation path for AI validation.** The system currently has a kill switch for PR secret scanning, but does not have a graceful degradation mechanism for the AI validation step. We are working on a mechanism that can disable AI model calls specifically for PR scans, allowing scans to continue while avoiding the dependency on a degraded model-serving layer. The goal is to allow scans to continue with reduced functionality rather than stopping PR secret scanning entirely.\n\n\u2022 **Improve alerting and incident communication.** The alerting infrastructure detected PR scan lag, but the status page was not updated promptly. We are improving the process for escalating service degradation and communicating customer-facing impact through the status page.\n\n\u2022 **Improve model-serving infrastructure resilience.** The underlying infrastructure issue occurred because nodes were scheduled on a subnet with no available IPs. The immediate issue was resolved by moving nodes to subnets with free IPs. We are continuing to improve the resilience of the model-serving infrastructure to reduce the likelihood that infrastructure capacity issues can affect PR scan processing.\n\n**Conclusion**\n\nThe incident was caused by degradation in the AI model validation service due to an infrastructure issue where nodes were scheduled on a subnet with exhausted IP capacity.\n\nThe resulting model validation timeouts slowed PR secret scan processing and caused consumer lag to build.\n\nWe mitigated the incident by temporarily stopping new PR secret scans, allowing the backlog to drain, and restoring model-serving capacity by moving nodes to subnets with available free IPs.\n\nWe are following up with improvements to model-serving resilience, timeout and retry behavior, runtime configuration, graceful degradation, alerting, and customer communication.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-08-14T16:14:19Z",
"resolved_inferred": false,
"started_at": "2026-08-14T13:41:27Z",
"state": "resolved",
"title": "Degraded performance in Secrets Pull Request scanning",
"updated_at": "2026-08-14T16:14:19Z"
},
{
"body": "GitHub component: Git Operations, Webhooks, Pull Requests\nOriginal GitHub incident: https://stspg.io/f0cgtdc0kqxg",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-08-13T15:39:18Z",
"resolved_inferred": false,
"started_at": "2026-08-13T14:53:48Z",
"state": "resolved",
"title": "GitHub incident can affect Cycode",
"updated_at": "2026-08-13T15:39:18Z"
},
{
"body": "GitHub component: Pull Requests\nOriginal GitHub incident: https://stspg.io/ssd9z8l2g46v",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-08-12T16:42:12Z",
"resolved_inferred": false,
"started_at": "2026-08-12T16:25:59Z",
"state": "resolved",
"title": "GitHub incident can affect Cycode",
"updated_at": "2026-08-12T16:42:12Z"
},
{
"body": "GitHub component: API Requests\nOriginal GitHub incident: https://stspg.io/3xn46bst0bjh",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-08-11T20:10:19Z",
"resolved_inferred": false,
"started_at": "2026-08-11T14:55:27Z",
"state": "resolved",
"title": "GitHub incident can affect Cycode",
"updated_at": "2026-08-11T20:10:19Z"
},
{
"body": "**Root Cause**\n\nThe incident was caused by a deployment of one service that ran a database index creation. The team has identified an issue with the way we perform index creations as database migrations. \n\nDuring a deployment the pods with newest image of the service attempted to create an index on a big table. The team has identified that the index creation took 7 minutes. However, during index creation pods were not responsive, and as a result, Kubernetes deemed them as unhealthy pods and attempted to retry those pods after 5 minutes.\n\nAs a result, because the pod got killed before the index creation was fully completed, the database transaction was rolled back. Then, subsequent pods attempted to create the index again, dying after 5 minutes.\n\nThis lead to the database being in unhealthy state, and the service was down. \n\nThe team has rolled back the deployment, and killed all replicas that attempted to create the index. Thanks to that, the service and the database was in healthy state again.\n\n \n\n**Why safety measures did not help**\n\nCycode provides a safety mechanism that unblocks all Pull Request scans after a specific period of time, giving each scan a maximum duration before the Pull Request is unblocked. However, because the service that is responsible for triggering and completing scans, as well as this safety net, was down, the process couldn't behave as expected. We acknowledge this gap and are working on strengthening this area of our system.\n\n \n\n**Action items**\n\n\u2022 The team is actively investigating enhancements and new safety protocols that can be put in place in order to have another safety net preventing Pull Request scans being stuck in case of any incident.\n\n\u2022 The team is investigating changes to the index creation process.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-08-10T15:45:54Z",
"resolved_inferred": false,
"started_at": "2026-08-10T14:06:34Z",
"state": "resolved",
"title": "We have noticed degraded performance in scanning",
"updated_at": "2026-08-11T15:59:42Z"
},
{
"body": "The platform is now fully operational and processing normally",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-08-04T06:08:42Z",
"resolved_inferred": false,
"started_at": "2026-08-02T14:08:56Z",
"state": "resolved",
"title": "We\u2019re investigating an issue causing older events to be reprocessed",
"updated_at": "2026-08-04T06:08:42Z"
},
{
"body": "**Root Cause**\n\nThe incident was caused by a deployment of one service that ran a database migration. The team has identified that this migration contained faulty code and as a result lead to database overload when attempting to deploy the service. As a result, the service was partially down until the deployment was reverted. During this time, all scans were processed with lower than expected performance.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-27T11:57:22Z",
"resolved_inferred": false,
"started_at": "2026-07-27T10:31:15Z",
"state": "resolved",
"title": "Degraded performance in scans",
"updated_at": "2026-07-31T07:43:19Z"
},
{
"body": "GitHub component: API Requests\nOriginal GitHub incident: https://stspg.io/vr201n49yl53",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-27T04:12:43Z",
"resolved_inferred": false,
"started_at": "2026-07-27T03:56:45Z",
"state": "resolved",
"title": "GitHub incident can affect Cycode",
"updated_at": "2026-07-27T04:12:43Z"
},
{
"body": "GitHub component: Pull Requests\nOriginal GitHub incident: https://stspg.io/sm1tp7kfm4vj",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-24T20:26:37Z",
"resolved_inferred": false,
"started_at": "2026-07-24T19:40:50Z",
"state": "resolved",
"title": "GitHub incident can affect Cycode",
"updated_at": "2026-07-24T20:26:37Z"
},
{
"body": "The issue has been resolved, and we are continuing to monitor the situation.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-24T21:58:42Z",
"resolved_inferred": false,
"started_at": "2026-07-24T17:45:35Z",
"state": "resolved",
"title": "Degraded performance in IaC Pull Request scans",
"updated_at": "2026-07-24T21:58:42Z"
},
{
"body": "GitHub component: API Requests, Pull Requests\nOriginal GitHub incident: https://stspg.io/j5c80shxqm53",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-24T17:40:25Z",
"resolved_inferred": false,
"started_at": "2026-07-24T16:21:08Z",
"state": "resolved",
"title": "GitHub incident can affect Cycode",
"updated_at": "2026-07-24T17:40:25Z"
},
{
"body": "GitHub component: Webhooks, Pull Requests\nOriginal GitHub incident: https://stspg.io/syhr80rth84z",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-23T09:39:23Z",
"resolved_inferred": false,
"started_at": "2026-07-23T07:58:43Z",
"state": "resolved",
"title": "GitHub incident can affect Cycode",
"updated_at": "2026-07-23T09:39:23Z"
},
{
"body": "The incident has been resolved, and the platform is operating normally.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-16T20:05:34Z",
"resolved_inferred": false,
"started_at": "2026-07-16T18:56:47Z",
"state": "resolved",
"title": "Platform Slowness",
"updated_at": "2026-07-16T20:05:34Z"
},
{
"body": "A fix has been applied and functionality is fully restored; we are continuing to monitor to ensure everything remains stable.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-15T19:07:05Z",
"resolved_inferred": false,
"started_at": "2026-07-15T17:22:05Z",
"state": "resolved",
"title": "Cycode CLI 3.17.1/2 Secrets scans failures",
"updated_at": "2026-07-15T19:07:05Z"
},
{
"body": "We\u2019re continuing to monitor detection processing and the associated delays in violation status updates (including auto-resolution)",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-14T21:23:08Z",
"resolved_inferred": false,
"started_at": "2026-07-14T16:12:50Z",
"state": "resolved",
"title": "Delays in violations status updates in EU",
"updated_at": "2026-07-14T21:23:08Z"
},
{
"body": "The system is now fully operational. There should be no more degraded performance.\n\n \n\n**Summary**\nWe observed a period of slowness and intermittent timeouts affecting various system functions in the EU region, including the application interface and pull request (PR) scans. The issue was primarily caused by a processing system reaching its network and memory capacity limits, exacerbated by a high volume of automated activity from a single source. We have since upgraded the underlying infrastructure and implemented safeguards to prevent similar high-volume activity from impacting the system. The issue is now fully resolved, and all services have returned to expected performance levels.\n\n**Key Timeline (IDT)**\n\n\u2022 **July 13, 2026, 11:44 IDT**: Incident detected following reports of UI slowness and PR scan delays.\n\n\u2022 **July 13, 2026, 12:19 IDT**: Infrastructure bottleneck identified; decision made to upgrade the processing cluster.\n\n\u2022 **July 13, 2026, 12:26 IDT**: A high-volume automated process was identified and disabled to reduce immediate load.\n\n\u2022 **July 13, 2026, 13:09 IDT**: Infrastructure upgrade completed; network throughput returned to normal levels.\n\n\u2022 **July 13, 2026, 15:35 IDT**: All backlogs cleared, and the incident was officially resolved.\n\n**Root Cause**\nThe incident was triggered by a combination of factors: a processing cluster reached its maximum network bandwidth and memory capacity due to an undersized configuration for the current workload. This was further strained by a specific automated workflow that generated an unusually high volume of update requests. Additionally, a configuration difference in the message processing pipeline in the EU region prevented the system from effectively handling the resulting backlog.\n\n**Actions Taken**\n\n\u2022 **Upgraded Infrastructure**: The processing cluster was upgraded to a higher-capacity instance type to provide more network bandwidth and memory.\n\n\u2022 **Disabled High-Volume Source**: A specific client identifier responsible for excessive traffic was temporarily disabled to restore system stability.\n\n\u2022 **Restored Connectivity**: Affected service components were restarted to ensure they re-established clean connections to the upgraded infrastructure.\n\n\u2022 **Increased Processing Parallelism**: The number of partitions in the affected message queue was increased to allow the system to process the backlog more quickly.\n\n**Action Items**\n\n\u2022 **Enhance Monitoring**: Implement new alerts for network and memory utilization to detect capacity issues before they impact customers.\n\n\u2022 **Optimize Update Workflow**: Refactor the status update process to batch requests, significantly reducing the load on the processing system.\n\n\u2022 **Implement Rate Limiting**: Introduce safeguards to prevent any single source from consuming disproportionate system resources.\n\n\u2022 **Standardize Regional Configurations**: Conduct an audit to ensure infrastructure and message queue settings are consistent across all regions.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-13T12:36:21Z",
"resolved_inferred": false,
"started_at": "2026-07-13T08:47:48Z",
"state": "resolved",
"title": "Degraded performance",
"updated_at": "2026-07-13T12:36:21Z"
},
{
"body": "Functionality is fully restored; we are continuing to monitor to ensure everything remains stable.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-12T12:30:15Z",
"resolved_at": "2026-07-17T16:58:39Z",
"resolved_inferred": false,
"started_at": "2026-07-09T10:54:26Z",
"state": "resolved",
"title": "Inaccurate violation counts in some dashboards and panels.",
"updated_at": "2026-07-17T16:58:39Z"
},
{
"body": "The issue affecting Maestro AI services has been resolved. Maestro Risk Explorability, Risk AI Remediation, Maestro Remediation, and Graph AI are now available and operating normally.\n\n \n\n**Summary**\n\nOn July 9, 2026, customers using the Maestro service in the European production environment experienced a period of service unavailability. The issue began following a configuration update that inadvertently changed the service's regional routing. This caused the system to attempt connections through a network path that lacked the necessary permissions and to a region where specific processing models were unavailable. The issue has been fully resolved, and service has been restored to all affected customers.\n\n**Key Timeline (IDT)**\n\n\u2022 **July 9, 2026, 12:02 IDT:** The incident was identified and an investigation was initiated.\n\n\u2022 **July 9, 2026, 12:07 IDT:** Public notification was issued regarding the service interruption.\n\n\u2022 **July 9, 2026, 13:00 IDT:** A network configuration fix was applied, restoring primary connectivity.\n\n\u2022 **July 9, 2026, 13:39 IDT:** Service was fully restored after implementing model fallbacks, and the incident was marked as resolved.\n\n**Root Cause**\n\nThe service interruption was triggered by a recent update to the authentication and configuration process. This update introduced a conflict in how the system identified its operating region. Specifically, an automated update process overrode manual settings, routing traffic to a different regional endpoint. This new path was blocked by a missing network security rule and attempted to use a processing model that was not supported in that specific region, leading to service failure.\n\n**Actions Taken**\n\n\u2022 **Restored Network Connectivity:** Manually updated network security rules to allow secure traffic through the new regional endpoint.\n\n\u2022 **Implemented Model Fallbacks:** Configured the system to use alternative processing models to ensure immediate service availability while long-term regional configurations were adjusted.\n\n\u2022 **Updated Status Communications:** Maintained real-time updates for stakeholders and customers throughout the recovery process.\n\n**Action Items**\n\n\u2022 **Standardize Configuration Precedence:** Update the deployment workflow to prevent automated processes from silently overriding critical environment settings.\n\n\u2022 **Infrastructure Audit:** Conduct a comprehensive review of network security rules across all regions to ensure consistency and prevent similar connectivity gaps.\n\n\u2022 **Enhance Automated Monitoring:** Implement end-to-end health checks and synthetic probes to detect regional connectivity issues automatically before they impact users.\n\n\u2022 **Improve Deployment Policies:** Establish new guidelines to ensure that configuration changes are deployed and validated in production-like environments more frequently to reduce the risk of \"stale\" updates.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-08T12:29:58Z",
"resolved_at": "2026-07-09T10:39:50Z",
"resolved_inferred": false,
"started_at": "2026-07-09T09:07:18Z",
"state": "resolved",
"title": "Maestro AI service disruption",
"updated_at": "2026-07-09T10:39:50Z"
},
{
"body": "AWS has confirmed that the issue has been fully mitigated and we are currently not observing any related issues.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-07T12:26:33Z",
"resolved_at": "2026-07-06T15:08:14Z",
"resolved_inferred": false,
"started_at": "2026-07-06T12:52:03Z",
"state": "resolved",
"title": "Infrastructure Provisioning Delays",
"updated_at": "2026-07-06T15:08:14Z"
}
]
}