{
"vendor": "Harness",
"slug": "harness",
"platform": "statuspage",
"status_url": "https://status.harness.io",
"last_checked": "2026-09-16T12:28:20Z",
"last_state": "degraded",
"history_backfilled": true,
"first_watched": "2026-09-04T07:06:16Z",
"incidents": [
{
"body": "Fix has been implemented in Prod Eu 1 and we are monitoring for any further issues.",
"first_seen": "2026-09-16T12:28:20Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"started_at": "2026-09-15T21:13:38.999-07:00",
"state": "monitoring",
"title": "IACM Module Registry timeouts",
"updated_at": "2026-09-16T00:00:47.701-07:00",
"url": "https://stspg.io/08kzqvspqx8l"
},
{
"body": "This incident has been resolved.",
"first_seen": "2026-09-16T12:28:20Z",
"impact": "critical",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-09-15T10:43:21.281-07:00",
"resolved_inferred": false,
"started_at": "2026-09-15T10:22:37.000-07:00",
"state": "resolved",
"title": "API Login failures",
"updated_at": "2026-09-15T12:28:12.095-07:00",
"url": "https://stspg.io/ft7mmvv8qjn8"
},
{
"body": "This incident has been resolved.",
"first_seen": "2026-09-15T12:28:18Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-09-15T07:58:13.180-07:00",
"resolved_inferred": false,
"started_at": "2026-09-15T05:05:00.982-07:00",
"state": "resolved",
"title": "IACM Module Registry performance degradation",
"updated_at": "2026-09-15T07:58:13.196-07:00",
"url": "https://stspg.io/r7lct0mc45rz"
},
{
"body": "This incident has been resolved.",
"first_seen": "2026-09-15T12:28:18Z",
"impact": "critical",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-09-14T20:28:11.132-07:00",
"resolved_inferred": false,
"started_at": "2026-09-14T12:00:28.000-07:00",
"state": "resolved",
"title": "MTLS Outage with Delegates",
"updated_at": "2026-09-14T20:29:00.938-07:00",
"url": "https://stspg.io/6mmv855y29p7"
},
{
"body": "### Summary\n\nOn September 8, 2026, customers in Prod 1 and Prod 2 experienced elevated platform latency and pipeline failures. The issue was caused by a regression in a newly released capability that triggered cascading failures under high load. Because the capability was behind a feature flag, it was quickly disabled, and service was restored after a brief monitoring period.\u00a0\n\n### Customer Impact\n\n* Customers encountered slowness and failures during pipeline execution and UI operations. Some API calls returned errors or timed out.\n* No data loss or corruption occurred.\n\n### Root Cause\n\nThe new capability introduced a regression that created contention on a shared backend resource used by multiple Harness components. This saturated the shared platform infrastructure and caused the cascading failures.\n\n### Mitigation\n\n* Disabled the capability across all environments\n* Temporarily increased platform capacity to restore stability\n\n### Next Steps\n\nTo prevent recurrence, Harness will:\n\n1. **Permanently fix the capability** by profiling and eliminating the sub-optimal code path and query\n2. **Improve detection** by enhancing alerting for resource-intensive queries on high-frequency platform paths",
"first_seen": "2026-09-09T12:30:34Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-09-08T21:30:49.347-07:00",
"resolved_inferred": false,
"started_at": "2026-09-08T14:52:48.883-07:00",
"state": "postmortem",
"title": "Slowness in Prod1 and Prod2 environment",
"updated_at": "2026-09-09T16:23:46.137-07:00",
"url": "https://stspg.io/nq6pjsvq1481"
},
{
"body": "## **Summary**\n\nBetween 12:34am PST and 12:38am PST on 5th September, the Delegate service manager experienced some elevated exceptions when attempting to write to the database. Consequently, delegate connections were dropped, causing them to disconnect. Delegate automatically re-attempts registration back to the `delegate service manager`\u00a0 and majority of the delegates got connected back after the incident. For Docker and ECS delegates the automatic restart is not enabled unless these delegates have health monitoring enabled. For these delegates a manual restart is needed and was recommended.\u00a0 Post restart the delegate would re-connect and the issue was resolved.\n\n## **Root cause**\n\nOn Prod2 cluster we identified a performance bottleneck in the delegate service that, under certain conditions, can increase database write latency and delay heartbeat processing which leads to delegates being disconnected.\n\n## **Impact**\n\nAll K8s delegates and \\`Docker/ECS\\` delegates got connected back immediately within 4 mins and started to function normally. The impact can be scoped to those specific types of delegates that didn\u2019t have health monitoring enabled.\u00a0\n\n## **Remediation**\n\n* Immediate: We have added additional monitoring and increased resources for handling the influx of traffic.\u00a0\n* Permanent: We have identified a hotspot in the code that can cause high latency when writing to a database which we are actively working on resolving.\u00a0\n\n## **Action Items**\n\nTo prevent such issues from happening again, Harness will work on the following:\n\n1. Increased targeted monitoring and alerting to initiate timely mitigation and prevent this from happening again.\n2. Fix the identified delegate service managers database client reconnect failures\n3. Fix the hotpots that can cause query latency.",
"first_seen": "2026-09-05T12:15:31Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-09-05T02:40:48.672-07:00",
"resolved_inferred": false,
"started_at": "2026-09-05T02:40:02.542-07:00",
"state": "postmortem",
"title": "Prod2 was intermittently unavailable",
"updated_at": "2026-09-09T11:22:04.182-07:00",
"url": "https://stspg.io/wxvc3v9g5yc2"
},
{
"body": "Proactive Service Scaling Notice: \nAs part of our ongoing efforts to support continued growth and ensure the scalability and reliability of our service, we will be scaling our production services during the scheduled maintenance window.\nThis activity involves making planned capacity adjustments across our production infrastructure. As the changes may span multiple clusters, customers may notice brief variations in service performance during the activity.\nOur engineering team will be closely monitoring the environment throughout the process to ensure a smooth and stable rollout.\nNo action is required from customers. This is a proactive infrastructure activity to ensure we maintain the capacity and reliability needed to support our growing usage.\nWe appreciate your understanding as we continue to improve and scale the platform.",
"first_seen": "2026-09-05T08:54:55Z",
"impact": "maintenance",
"last_seen": "2026-09-05T08:54:55Z",
"resolved_at": "2026-09-05T12:15:31Z",
"resolved_inferred": true,
"started_at": "2026-09-04T22:30:00.000-07:00",
"state": "maintenance",
"title": "CI Cloud Maintenance",
"updated_at": "2026-09-04T22:29:36.680-07:00",
"url": "https://stspg.io/8dtjn0v78mq0"
},
{
"body": "# Summary\n\nOn September 3, 2026, between approximately 2:54 AM and 3:44 AM PDT, customers running pipelines on Prod1 experienced delays in pipeline execution graph rendering and execution status updates. Pipeline execution itself was not affected \u2014 pipelines continued to run and complete \u2014 but the visual graph and status information in the Harness UI lagged behind actual execution progress. Harness engineering identified the cause, added processing capacity, and restored normal graph and status updates. Extended monitoring confirmed full recovery the following day.\n\n# Root Cause\n\nA planned database failover activity in the Prod1 region temporarily increased network latency between the pipeline execution service and its database while traffic was briefly served cross-region. This reduced the rate at which the service could process pipeline orchestration events, and a processing backlog began to build.\n\nPipeline execution graphs and status indicators in the UI depend on these orchestration events being processed in near real time. As the backlog grew, graph rendering and status updates fell increasingly behind actual execution progress, and some graphs had to be rebuilt rather than served from cache. Pipeline execution itself does not depend on this same processing path, so pipelines continued to run and complete throughout the incident.\n\n# Impact\n\n* Affected users: Customers with pipeline executions running on Prod1 during the incident window.\n* Symptom: Delayed pipeline execution graph rendering and delayed execution status updates in the Harness UI.\n* Not reported as impacted: Pipeline execution itself \u2014 pipelines continued to run and complete.\n* Data: No data loss, corruption, or exposure was identified.\n* Suggested workaround: None required from customers. The issue was resolved entirely by Harness engineering.\n* Duration of impact: Approximately 45\u201350 minutes of degraded graph and status visibility, from 2:54 AM to 3:44 AM PDT on September 3, 2026. Harness continued monitoring after the backlog cleared and confirmed full recovery.\n\n# Remediation\n\n## Immediate\n\nHarness engineering added processing capacity to the affected service and restarted it, which cleared the event-processing backlog and restored normal pipeline execution graph and status updates.\n\n# Action Items\n\n| **Action Item** | **Objective** |\n| --- | --- |\n| Queue based alarms | We are raising priority for queue based alarms so we can take faster remediation actions \\(increase capacity\\) |",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-09-03T04:40:52.958-07:00",
"resolved_inferred": false,
"started_at": "2026-09-03T03:00:57.335-07:00",
"state": "postmortem",
"title": "Pipelines are stuck in Prod1",
"updated_at": "2026-09-10T00:26:17.980-07:00",
"url": "https://stspg.io/1l82qknlzdl6"
},
{
"body": "## Summary\n\nBetween August 27 and August 28, 2026, customers experienced an issue where some pipelines, deployments, and related resources appeared as not found in the Harness UI and API, even though the underlying data remained intact.\n\n\u200c\n\nThe issue occurred during a planned internal infrastructure update that affected communication between internal platform services. As a result, requests that depended on account, organization and project scope resolution were unable to complete successfully, which led to incorrect not found responses being returned to customers for existing entities.\n\n\u200c\n\nEngineering identified the issue, rolled back the change, and restored normal service. No customer data was lost or deleted during the incident.\n\n\u200c\n\n## Root Cause\n\nThe issue was caused by a configuration error introduced during a planned internal service routing update in Production.\n\nAn internal platform service responsible for resolving account, organization, and project context was unable to validate requests from other Harness services after the change was applied. Because that validation step is required before many entity reads and pipeline-related actions can proceed, the failed requests surfaced to customers as not found errors for resources that continued to exist normally.\n\nThe issue was limited to the affected production environment and was resolved by reverting the change and restoring the previous service communication path.\n\n\u200c\n\n## Impact\n\n* Some customers saw existing pipelines, deployments, and related entities appear as not found in the UI and API.\n* Some pipeline-related operations, including execution progression, webhook-triggered starts, scheduled trigger evaluation, and entity listing, were temporarily disrupted.\n* The issue affected availability and visibility of existing entities, but it did not remove data or change customer configurations.\n* No unauthorized access occurred, and no customer data loss was observed.\n\n## Remediation\n\n* **Immediate:** Reverted the infrastructure configuration update and restored the previously working service communication path.\n* **Recovery validation:** Verified that affected entity lookups, pipeline operations, and dependent APIs were functioning normally after rollback.\n* **Permanent:** Corrected the configuration handling associated with the update so similar issues do not interfere with service-to-service authentication in future rollouts.\n\n## Action Items\n\nTo prevent such issues from happening again, Harness will\u00a0\n\n1. Improve configuration validation by enhancing the pre-deployment tests to verify internal service communication before shifting production traffic.\n2. Enhance monitoring and alerting for internal authentication failures so issues can be detected earlier.\n3. Improve error handling so dependency failures are less likely to appear to customers as resource not found errors.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-08-28T00:35:45.213-07:00",
"resolved_inferred": false,
"started_at": "2026-08-28T00:04:09.101-07:00",
"state": "postmortem",
"title": "Entities in Harness are not loading on in Prod3",
"updated_at": "2026-09-02T15:16:13.803-07:00",
"url": "https://stspg.io/y4dw5tg21696"
},
{
"body": "This incident has been resolved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-08-26T02:10:35.577-07:00",
"resolved_inferred": false,
"started_at": "2026-08-26T01:38:48.000-07:00",
"state": "resolved",
"title": "Pipelines are failing for harness IACM customers",
"updated_at": "2026-09-02T15:16:34.604-07:00",
"url": "https://stspg.io/rgzrn5j4w820"
},
{
"body": "## Summary\n\n* Starting at **23:42 UTC** on August 23, 2026, several FME customers reported failures loading the FME UI.\n* FME UI Artifacts served from the CDN expired due to a retention policy, causing FME UI to fail to load.\n* Any flag request changes through the API, change delivery, and the data pipeline continued to work with no interruption.\n\n## Root Cause\n\n* The FME UI is served from a CDN. The UI artifacts got evicted due to a retention policy, causing the UI to fail to load for all users.\n\n## Impact\n\n* The FME UI was unable to load for all users across all production environments.\n\n### What was not impacted?\n\n* SDK functionality and runtime flag evaluation\n* Admin API calls\n* Customer flag configuration data\n* No data loss occurred\n\n## Remediation\n\n* FME UI got restored in the CDN through a deployment\n* Recovery confirmed across all production environments before closing the incident.\n\n## Action Items\n\n* Improve the asset retention policy so that the currently active version is never subject to eviction.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-08-23T18:12:48.700-07:00",
"resolved_inferred": false,
"started_at": "2026-08-23T17:57:07.808-07:00",
"state": "postmortem",
"title": "Feature Management & Experimentation (FME) user interface unavailable",
"updated_at": "2026-08-27T16:37:23.848-07:00",
"url": "https://stspg.io/jf7z6c8ycf22"
},
{
"body": "# Summary\n\nOn 20 August 2026, beginning at approximately 15:00 UTC, the Harness platform experienced widespread performance degradation across all production environments. Pipeline executions that normally complete in around two minutes took seven to ten minutes. Continuous Delivery, Continuous Integration, pipeline orchestration, and Feature Management & Experimentation were all affected.\n\nGoogle Cloud Platform experienced a multi-product incident in the us-west1 region affecting Bigtable, Compute Engine, Google Kubernetes Engine, and persistent-disk I/O. Harness production infrastructure runs on persistent disks in that region. The degradation raised database operation latency from approximately 2 ms to over 10 ms at the 95th percentile, which in turn caused message-queue processing lag and propagated to every service that depends on timely database access.\n\n\u200c\n\n# Impact\n\nThis was a degradation, not an outage. Pipelines continued to execute and complete successfully throughout; they were slow rather than failing. No data was lost, and no customer work was dropped as a result of this incident.\n\n# **Root cause**\n\nHarness production infrastructure in the affected environments runs on Google Cloud Platform persistent disks in the us-west1 region. When that storage layer degraded, the effect propagated through the platform in a predictable chain:\n\n**Persistent-disk I/O degradation in us-west1.** Google Cloud Platform experienced a multi-product incident affecting Bigtable, Compute Engine, Google Kubernetes Engine, and persistent-disk performance. This was an infrastructure failure in the provider\u2019s environment, outside Harness\u2019s control.\n\n# **Preventive actions**\n\nAlthough Harness cannot prevent a cloud provider infrastructure failure. The actions below are aimed at detecting one faster and being better positioned to act on it.\n\n| **Action** |\n| --- |\n| Continue routine pre-testing of targeted cross-region database failovers, as performed during this incident, to keep failover readiness verified rather than assumed |\n| Assess full-stack multi-region failover readiness for future scenarios in which cross-region latency would be unacceptable |",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-08-20T12:23:42.742-07:00",
"resolved_inferred": false,
"started_at": "2026-08-20T08:37:35.998-07:00",
"state": "postmortem",
"title": "All modules are running slow in Prod1/2/3/4 due to cloud provider incident",
"updated_at": "2026-08-26T09:20:43.140-07:00",
"url": "https://stspg.io/jn63dxz0f5ty"
},
{
"body": "### Summary\n\nOn August 20, 2026, between 10:24 and 14:55 UTC, a subset of FME writes failed. Writes made from the FME UI and writes made with Harness access tokens \\(PATs and SATs\\) were not affected. Runtime flag evaluation continued to work normally. The issue was mitigated by reverting a recent authentication change in a shared governance service, and affected writes returned to normal by 14:55 UTC. Status: [https://status.harness.io/incidents/rhthgm7d5dkz](https://status.harness.io/incidents/rhthgm7d5dkz)\n\n### Root Cause\n\nA change in how a shared governance service authenticated inbound calls resulted in some FME writes being rejected. Those writes used service-to-service credentials that the governance service could no longer verify after the change. FME surfaces a governance failure to the client as HTTP 499,  the same status used when a governance policy intentionally denies a change. Because 499 is a valid, expected response in that deny path, the failures did not look like an outage on our alerts, and the incident was identified from customer reports rather than internal detection.\n\n### Impact\n\n* A subset of FME writes failed during the window, primarily those made using legacy Split API keys or change request scheduling.\n* Writes made from the FME UI were not impacted.\n* Writes using Harness access tokens \\(PATs and SATs\\) were not impacted.\n* Runtime flag evaluation continued normally.\n* No data loss occurred. Failed writes did not apply.\n\n\u200c\n\n### Remediation\n\nReverted the governance-service authentication change. Affected writes returned to normal immediately.\n\n### Action Items\n\nTo prevent such issues from happening again, \n\n* Harness will return a distinct error \\(not 499\\) when a write fails because governance could not be evaluated, so it is not confused with an intentional policy denial.\n* Add alerting on the governance evaluation call itself, rather than relying on the client-facing status code.\n* Expand authentication support for policy evaluations.\n* Expand automated coverage for additional write scenarios.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-08-20T08:02:26.071-07:00",
"resolved_inferred": false,
"started_at": "2026-08-20T07:32:40.483-07:00",
"state": "postmortem",
"title": "FME API write operations started returning 499 errors",
"updated_at": "2026-08-23T14:46:49.187-07:00",
"url": "https://stspg.io/wnr34j93m8rm"
},
{
"body": "**Summary**\n\nOn 19 August 2026 between 12:35 and 17:29 UTC, the Harness Application Security service experienced a significant disruption affecting both the customer-facing console and the data ingestion pipeline in the SaaS Production and US1 regions.\n\n\u200c\n\n**Root Cause** \n\nThe internal configuration service that supplies runtime settings to nearly every other component became overloaded and entered a repeated restart cycle. Because so many services depend on it, the effects were broad: console pages such as protection policies, posture views, activity logs, API inventory, and custom policy failed to load or timed out, and downstream processing stalled while waiting for configuration it could not obtain.\n\n# **Customer impact**\n\n| **Dimension** | **Detail** |\n| --- | --- |\n| Console \\(UI\\)  impact | Multiple pages failed to load or timed out, including protection policies, posture event pages and posture views inside dashboards and insight pages, activity log queries, API inventory screens, custom policy, and sensitive-data views and widgets. |\n| Ingestion impact | Security telemetry processing degraded severely and, in some paths, stopped entirely. Consumer lag grew across normalisation, grouping, anomaly detection, generation, and related processing stages.  |\n| Data loss | A subset of telemetry ingested during the disruption was permanently dropped. |\n\n\u200c\n\n**Mitigation**\n\n Several intermediate mitigations  additional CPU and memory, relaxed health-check thresholds, a database restart, and a larger connection pool ameliorated the issue. Disabling the new feature in both affected regions restored throughput sharply and durably. The incident was resolved at 17:29 UTC.\n\n\u200c\n\n# **Preventive actions**\n\nThe following actions are committed and tracked internally to completion. The feature that triggered this incident remains disabled and will not be re-enabled until the work below is complete and validated.\n\n| **Action** |\n| --- |\n|  |\n| OPtimize the code by tuning parameters such as  cache eviction and retention , evaluate cursor-based pagination for bulk rule retrieval as rule counts grow |\n| Add a purpose-built database index for the service-scoping access pattern |\n| Remediate pipeline recovery semantics so consumers replay safely after position-marker loss instead of skipping backlog |\n| Mandate staged rollout for configuration overrides that alter downstream request patterns: low-volume cluster, then mid-volume, then high-volume |\n| Add backpressure and concurrency protection to the configuration service: circuit breaking, bounded queues, and timeout isolation |\n| Enhance observability by Instrumenting more detailed metrics  |",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-08-19T09:07:13.081-07:00",
"resolved_inferred": false,
"started_at": "2026-08-19T06:22:48.975-07:00",
"state": "postmortem",
"title": "Data ingestion is delayed on Traceable US production",
"updated_at": "2026-08-27T16:35:38.537-07:00",
"url": "https://stspg.io/l20k0l87d9qy"
},
{
"body": "## Summary\n\nCustomers on Prod1, Prod2, and Prod3 \\(US\\) clusters experienced failures when loading SEI 2.0 dashboards on August 6, 2026, from 7:22 AM PDT to 9:03 AM PDT. Customers calling the SEI 2.0 API also experienced similar failures.\n\nNo customer data was lost, and ingestion of all integration data continued to work uninterrupted. SEI customers using 1.0 were not impacted.\n\n## Root Cause\n\nThe incident was caused by resource exhaustion on the nodes serving queries. This resource degradation developed in a pattern that did not cross our existing alerting thresholds early enough to provide sufficient warning or allow mitigation before customer impact occurred.\n\n## Impact\n\nCustomers on Prod1, Prod2, and Prod3 \\(US\\) clusters were unable to load SEI 2.0 dashboards during the incident window.\n\n**Duration:** August 6, 2026, 7:22 AM PDT \u2013 9:03 AM PDT \\(~1 hour 41 minutes\\)\n\n### What was not impacted?\n\n* Data ingestion and processing\n* SEI 1.0 customers\n* Integrations and metadata flows\n\nNo customer data was lost.\n\n## Remediation\n\nUpon identifying the root cause, our team took immediate corrective action by adding capacity to restore the affected systems. Services were fully recovered, and all dashboards resumed normal operation at 9:03 AM PDT.\n\n## Action Items\n\nTo prevent from such issues happening again, Harness is/has\n\nProactively added  capacity  updates have been applied to prevent this issue from recurring\n\n####  Enhanced Monitoring and Alerting\n\nAdditional monitoring and alerting have been put in place to detect anomalies early, focused on a leading indicator, which in this case was thread pool exhaustion, before they can impact dashboard availability and data rendering.\n\n#### System Patch in Progress\n\nWe are working with our vendor to apply a patch to remediate this and similar issues completely.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-08-06T09:01:54.237-07:00",
"resolved_inferred": false,
"started_at": "2026-08-06T07:30:56.391-07:00",
"state": "postmortem",
"title": "SEI 2.0 dashboards are not loading",
"updated_at": "2026-08-13T21:17:19.148-07:00",
"url": "https://stspg.io/twn4gg4bm5mz"
},
{
"body": "## **Summary**\n\nOn August 6, 2026 \\(morning PDT\\), some customers running pipelines in the Prod2 production environment observed pipeline executions that stopped making progress \u2014 stages that did not advance and produced no further output or status updates. The issue was reported by affected customers. Harness engineers identified the cause, mitigated the impact, and pipeline executions returned to normal operation.\n\nThe issue was caused by a self-referential pipeline expression. A Git webhook triggered a pipeline that referenced the contents of the webhook payload, and the payload itself contained further copies of that same expression. Each round of expression resolution therefore produced more expressions to resolve, doubling the amount of work each time. This exhausted the resources of the service instance processing that execution, and other executions assigned to the same instance were unable to progress while it was in that state.\n\n## **Impact**\n\nDuring the incident window \\(approximately 6:11 AM to 11:23 AM PDT on August 6, 2026\\):\n\n* Some customers' pipeline executions on Prod2 stalled mid-execution and made no further progress.\n* Affected executions produced no new step output or status updates, and had to be aborted and re-run after mitigation.\n* Behavior was limited to executions being processed by the affected service instance \u2014 pipelines handled by other instances continued to execute normally.\n\nThere was **no data loss**. Pipeline definitions, execution history, and stored state were unaffected. The majority of pipelines on Prod2 continued to execute successfully throughout the incident; the primary impact was that some in-flight executions could not complete and needed to be re-run once the issue was mitigated.\n\n## **Root Cause**\n\nHarness pipelines support expressions that are resolved at runtime \u2014 for example, an expression that inserts the contents of the Git webhook payload that triggered the pipeline.\n\nIn this case, a Git commit message contained the literal text of the payload expression itself, twice, and the pipeline referenced that same payload expression. Because the commit message is part of the webhook payload, resolving the expression inserted the entire payload \u2014 including the two literal copies of the expression carried in the commit message. Those newly inserted copies were then treated as expressions to be resolved, and each pass inserted two more full copies of the payload. The size of the value being processed, and the work required to process it, therefore doubled on every pass and grew exponentially rather than converging.\n\nHarness has a safeguard intended to stop exactly this: expression resolution is bounded by a maximum nesting depth, beyond which resolution halts and the pipeline fails with an explicit error. A defect in that safeguard meant the limit was not applied in this specific self-referential case, so resolution continued unchecked.\n\nExpression resolution runs inline on the threads that start pipeline steps. As each pass consumed progressively more memory and CPU without ever completing, the service instance performing that work stopped making progress, and every execution assigned to that instance stalled \u2014 which is what customers reported.\n\n## **Mitigation**\n\nHarness completed the following immediate mitigation steps:\n\n* Identified the pipeline and the expression pattern responsible for the runaway resolution.\n* Stopped the affected service instance so that it would take on no further work. The remaining healthy instances picked up and processed queued executions normally.\n* Confirmed that pipeline executions returned to normal and closed the incident.\n\nThese actions restored normal pipeline execution behavior and resolved the customer-facing impact.\n\n## **Action Items**\n\nTo reduce the risk of recurrence and improve detection, the following actions are in various stages of being implemented:\n\n* Fix the defect in the expression depth and loop-detection safeguard so that self-referential expressions are caught and fail fast with a clear error instead of consuming resources without bound.\n* Prevent payload expressions from being resolved out of trigger payload content, removing the self-referential path entirely.\n* Tighten the maximum expression nesting depth and evaluate explicit loop detection in addition to the existing depth limit.\n* Enhance automated tests in pre-production environments that reproduce self-referential expression patterns and verify that the safeguard detects and stops them.\n* Add monitoring for this pattern in pipeline executions so that it is detected proactively.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-08-06T11:21:32.289-07:00",
"resolved_inferred": false,
"started_at": "2026-08-06T06:50:18.466-07:00",
"state": "postmortem",
"title": "Monitoring - Pipelines Stuck - Prod2",
"updated_at": "2026-08-17T11:42:16.472-07:00",
"url": "https://stspg.io/nzw13j14pl74"
},
{
"body": "# Executive Summary\n\nOn August 4, 2026, between approximately 3:36 PM and 9:00 PM IST, customers using Infrastructure as Code Management \\(IaCM\\) on Prod0 and Prod1 were unable to access the Variable Sets settings page. The page rendered blank with no error message, and customers with Variable Sets attached to their workspaces could not view or manage them for the duration of the incident. Prod2, Prod3, and EU1 were not affected.\n\nSeparately, during the same window, a scheduled maintenance action caused the IaCM settings tab to temporarily disappear across all environments. This was identified and reversed within the incident bridge call before significant customer impact occurred.\n\nWe deployed a hotfix that restored full access to the Variable Sets page on Prod0 and Prod1 the same evening, and we are implementing permanent safeguards described below to prevent this class of issue from recurring.\n\n# Impact\n\n* Customers with the Variable Sets feature enabled on Prod0 and Prod1 were unable to view or manage Variable Sets for approximately 5\u20136 hours.\n* No data was lost or corrupted,  this was a UI routing failure only; underlying Variable Sets data and configuration were not affected.\n* Prod2, Prod3, and EU1 were not affected by this issue.\n* A secondary issue, a scheduled feature flag operation  caused the IaCM settings tab to temporarily disappear across all environments during the incident bridge call. This was identified and reversed within minutes. External customer exposure for this secondary issue is still being confirmed.\n\n# Root Cause\n\nA platform routing change released on July 18, 2026 updated how IaCM settings pages are resolved in the user interface. As part of that change, any settings page that had not been explicitly re-registered in the new routing structure became unreachable.\n\nThe Variable Sets page had not been re-registered under the new routing structure, making it inaccessible in the environments where the routing change had been deployed \u2014 Prod0 and Prod1. Because the failure occurred at the routing layer rather than within the page itself, the page rendered blank with no visible error rather than showing a clear failure message.\n\n# Remediation\n\n## Immediate\n\nWe deployed a hotfix that re-registered the Variable Sets page in the updated routing structure, restoring access for all affected customers on Prod0 and Prod1.\n\n## Permanent\n\nWe are adding automated end-to-end tests that navigate to settings pages with relevant feature flags enabled, configured as a required gate in our release pipeline. We are also documenting and enforcing the routing constraint through static analysis so that settings pages are never inadvertently left out of the routing structure during future platform changes.\n\n#  Action Items \n\nTo prevent such issues from happening again, \n\n1. Enhance automated tests, that navigate to settings pages with relevant feature flags enabled, configured as a blocking gate in the release pipeline, so this class of regression is caught before it reaches production.\n2. Establish an explicit checklist step for future platform-wide architectural changes that verifies all existing settings pages remain accessible in the updated routing structure before the change is promoted to production.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-08-04T08:28:40.558-07:00",
"resolved_inferred": false,
"started_at": "2026-08-04T05:07:20.745-07:00",
"state": "postmortem",
"title": "Editing 'Variable Sets' in the IaCM module is experiencing issue",
"updated_at": "2026-08-06T10:07:11.242-07:00",
"url": "https://stspg.io/n87ctg0rb2bv"
},
{
"body": "# **Summary**\n\nBetween 25 July and 4 August 2026, pipeline execution dashboards and overview pages in the Harness Prod 2 and Prod 3 clusters displayed data that was between  behind real time. Pipelines themselves continued to build, deploy, and execute normally throughout; the issue was confined to how quickly execution records were copied into the database that serves reporting and dashboard views.\n\n\u200c\n\n**No customer data was lost.** Every affected record remained durably stored and was replayed into the analytics datastore once the underlying limitation was removed. Harness migrated the affected clusters to a horizontally scalable, queue-backed version of the replication component on 1 August 2026 and completed targeted data backfills for all affected accounts. \n\n# **Root cause**\n\nHarness maintains a change-data-capture component that continuously replicates pipeline execution records from the primary operational datastore into a separate time-series datastore optimised for dashboards and reporting queries. Dashboards read exclusively from the analytics datastore. When replication falls behind, dashboards render an accurate but older view of the world, while execution itself is unaffected. This was caused by sharp, sustained increase in database write volume from another Harness platform module sharing the same replication path exceeded the throughput ceiling of the older, single-instance version of that component still running in Prod 2 and Prod 3. A backlog formed and grew.\n\n\u200c\n\n\u200c\n\n# **Preventive actions**\n\nHarness has completed or committed to the following actions to prevent such issues.\n\n| **Action** |\n| --- |\n| Fine tune the replication lag alerting so that any delay beyond a defined threshold is notified  |\n| Add a replication lag panel to the standard platform monitoring board so pipeline health is visible to on-call by default |\n| Reduce write amplification from co-tenant modules through per-module rate limiting or entity filtering on the replication stream |",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-08-01T02:11:14.661-07:00",
"resolved_inferred": false,
"started_at": "2026-07-31T13:22:52.702-07:00",
"state": "postmortem",
"title": "UI dashboards are lagging behind (CI)",
"updated_at": "2026-08-20T21:23:37.252-07:00",
"url": "https://stspg.io/9wsfx9tx5dbl"
},
{
"body": "# **Summary**\n\nOn July 31, 2026, artifact uploads performed through pipeline in the EU1 cluster began failing with an authentication error. Uploads initiated manually \\(outside of a pipeline\\) were not affected, and the ability to retrieve existing artifacts \\(downloads\\) was also unaffected \u2014 this was isolated to the specific pipeline upload path in one cluster.\n\n# **Impact**\n\n* Artifact uploads performed through pipeline in the EU1 cluster failed with an authentication error for approximately 4 hours and 34 minutes.\n* Retrieving existing artifacts \\(downloads\\) was not affected.\n* Manually uploading artifacts outside of a pipeline was not affected.\n* Other clusters/regions were not affected by this issue.\n\n# **Root Cause** \n\nThe component responsible for handling pipeline-based artifact uploads is distributed as a container image. In the EU1 cluster, this image is retrieved from an internal registry that mirrors a public image source; in other clusters, the same image is retrieved directly from the public source.\n\nA publishing error in our release process caused a new build of this component to be published using a version label that was already in use, rather than being assigned a new, unique version. As a result, two different images ended up associated with the same version label in the public source.\n\nOur internal registry mirrors images from the public source via an automated replication process. Because of how that replication was triggered, it copied the original \\(earlier\\) image associated with that version label rather than the corrected one. This meant the EU1 cluster \u2014 which pulls from the internal mirror \u2014 ended up running a different, defective image than other clusters, which pull directly from the public source and therefore received the corrected image. The defective image contained an authentication issue that caused pipeline uploads to fail.\n\n# **Mitigation**\n\n* Reverted the affected account to the last known-good version of the upload component, immediately restoring pipeline uploads.\n* Published a corrected, permanent version of the component to resolve the issue across all clusters.\n\n# **Next steps** \n\n\u200c\n\n* Fix the upload step to remove the underlying container-related defect that made this failure mode possible.\n* Update our release pipeline for this component so that publishing an image can never overwrite an existing version \u2014 every publish must create a new, distinct version going forward.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-08-01T16:49:44.458-07:00",
"resolved_inferred": false,
"started_at": "2026-07-31T06:48:52.846-07:00",
"state": "postmortem",
"title": "Harness Artifact Registry upload is failing from pipeline - EU1 region",
"updated_at": "2026-08-12T17:54:23.414-07:00",
"url": "https://stspg.io/gc7mqtsqf2ck"
},
{
"body": "## Summary\n\nStarting on August 4, 2026, CI runners in the us-west1 and us-central1 regions intermittently experienced connection timeouts of approximately 134 seconds when reaching external services such as GitHub and Bitbucket over outbound network gateways.\n\n## Impact \n\n* CI runners in the affected regions intermittently experienced connection timeouts of approximately 134 seconds when reaching external services \\(e.g., GitHub, Bitbucket\\) over our outbound network gateways.\n* The issue was intermittent rather than constant \u2014 connections succeeded under normal load, and failures clustered during periods of high outbound traffic volume.\n* No data was lost or corrupted. This was a network-connectivity and capacity issue, not a data-integrity issue.\n* us-west1 and us-central1 were the affected regions; other regions were not impacted by this issue.\n\n## Root Cause \n\n\u200c\n\nOur load balancer distributes outbound traffic across multiple NAT gateways using a hashing method based on connection details \\(source/destination address and port\\). For any single connection, these details stay constant for that connection's lifetime. We had a sustainted traffic surge for a few seconds which congested the gateways\n\n\u200c\n\n## Action Items\n\nTo prevent such issues from happening again Harness will,\n\nIncrease outbound connection capacity on our NAT gateways by provisioning additional external network interfaces, giving each gateway a substantially larger pool of connections it can serve concurrently..",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-30T20:22:54.084-07:00",
"resolved_inferred": false,
"started_at": "2026-07-29T22:59:09.783-07:00",
"state": "postmortem",
"title": "Intermittent External Network Connectivity Issues Affecting Build VMs",
"updated_at": "2026-08-07T13:24:14.611-07:00",
"url": "https://stspg.io/qdjcv8bw5w5x"
},
{
"body": "# Summary\n\nDuring a recent production deployment, a defect in our internal deployment tooling caused two critical services \u00a0 to run with incorrect, non-production configuration values in our production environment This led to a related set of four distinct symptoms: incorrect configuration behavior, intermittent login/access failures, a filestore access issue affecting one customer environment, and delayed pipeline status updates in the UI.\n\n\u200c\n\nWe have identified and are implementing a permanent fix for the underlying configuration defect, and have already put in place resource and capacity changes that resolve the UI delay symptom.\n\n\u200c\n\nAt no point during this incident were pipeline executions themselves lost, corrupted, or left in a stuck state. Where execution behavior was affected, it was limited to delays in status visibility, not in the underlying processing.\n\n# Incident Details\n\n## Incorrect Production Configuration Values Applied\n\nOur engineering team confirmed a defect in the Service Manager deployment pipeline that caused certain production services to be deployed using configuration values intended for a different environment, rather than the correct production configuration.\n\n\u200c\n\n**Root Cause**\n\n\u200c\n\nThe service responsible for fetching configuration overrides during deployment queries an internal API that returns a maximum of 1,000 results per request. The total number of services in the environment recently grew beyond that limit. As a result, any service beyond the first 1,000 returned was not included in the response, and the deployment pipeline silently fell back to default configuration values for those services. This is a confirmed pagination defect in the deployment tooling, not an issue with the configuration values themselves.\n\n\u200c\n\n**Resolution**\n\n\u200c\n\nEngineering has confirmed the mechanism and is implementing a permanent fix to remove this limit-related gap in the deployment pipeline.\n\n## Intermittent Login / Access Failures\n\nDuring the Service Manager deployment referenced above, some users experienced intermittent login or access failures. Under normal operation, previously running instances should continue serving traffic without interruption while a new deployment is in progress. In this incident, that fallback behavior did not occur as expected, contributing to access failures during the deployment window.\n\n## Filestore Access Issue\n\nA filestore access issue was identified that was specific to the Prod-3 environment and affected a single customer's environment.\n\n**Root Cause**\n\nThis is related to an IAM / storage-bucket permission configuration on Service Manager, potentially triggered by rollback activity.\u00a0\n\n## Delayed Pipeline Execution Status Updates in UI\n\nSome users observed that the pipeline execution graph in the UI was slow to refresh and did not reflect the latest status promptly. Importantly, this was a visibility delay only: there was no impact to actual pipeline executions, and no executions were stuck or failed as a result of this issue.\n\n**Root Cause**\n\nThe pipeline execution graph relies on a message stream \\(the orchestration log\\) to receive status updates. During the incident window, consumer processing of this stream fell behind \\(high consumer lag\\), which delayed how quickly status updates reached the UI. This was caused by the fact that the underlying database was in the middle of a planned scaling operation at the same time, and a traffic spike during that window further exacerbated the delay. Users experienced this as apparent pipeline slowness, even though the underlying executions were running normally.\n\n**Resolution**\n\nWe have increased resource capacity for the affected components to maintain more than 50% spare headroom going forward, reducing sensitivity to similar load spikes. This change has been implemented and is currently being validated as part of longer-term hardening for this part of the platform.\n\n# Impact Summary\n\n* Service Manager and License Manager ran with incorrect configuration values in the Prod-1 and Prod-3 environments.\n* Some users experienced intermittent login or access failures during the affected deployment window.\n* One customer environment in Prod-3 experienced a filestore access issue.\n* Users across affected environments saw delayed pipeline execution status updates in the UI; underlying pipeline executions continued to run correctly and were not lost, stuck, or corrupted.\n\n# Preventive Actions\n\nThe following corrective and preventive actions have been identified.\n\n\u200c\n\n| **Corrective / Preventive Action** |\n| --- |\n| Correct the pagination limit in the configuration-lookup service so that all services are returned and evaluated, regardless of total count. |\n| Add safeguards so that a service which cannot retrieve its configuration fails safely \\(e.g. alerts and blocks the deployment\\) rather than silently falling back to non-production defaults. |\n| Increase Postgres and messaging-pipeline resource headroom \\(target: greater than 50% spare capacity\\) to reduce sensitivity to concurrent load and scaling events. |\n\n\u200c\n\n_We recognize the impact this incident had across multiple areas of the platform and appreciate your patience as we work through a complete resolution._",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-27T06:33:56.030-07:00",
"resolved_inferred": false,
"started_at": "2026-07-26T23:42:31.644-07:00",
"state": "postmortem",
"title": "The Prod3 & Prod1 environment is experiencing intermittent outages. We are currently investigating the issue.",
"updated_at": "2026-08-07T11:16:09.379-07:00",
"url": "https://stspg.io/d5bcx0frs63p"
},
{
"body": "## Summary\n\nCustomers on Prod1, Prod2, and Prod3 clusters experienced intermittent widget load failures and increased load times when accessing AIDI 2.0 dashboards on July 22, 2026. Not all widgets were affected simultaneously  the issue manifested as sporadic failures rather than a full outage.\n\nNo customer data was lost. SEI 1.0 customers were not impacted.\n\n## Root Cause\n\nOver time, a routine database maintenance process failed to run on certain tables in our analytics database, causing those tables to accumulate a large volume of internal metadata used to track deleted records. When the database planned queries against these tables, it loaded all of this accumulated metadata into memory, causing memory usage on the affected nodes to spike repeatedly. These repeated spikes triggered an automatic safety mechanism that restarts a node when it detects excessive memory pressure, and the affected nodes began restarting in a loop as a result. This caused intermittent, degraded query performance on AIDI 2.0 dashboards for the duration of the incident.\n\n## Impact\n\nCustomers on Prod1, Prod2, and Prod3 clusters may have experienced intermittent widget load failures or increased load times on AIDI 2.0 dashboards.\n\n**Duration:** July 22, 2026, 07:58 PDT \u2013 16:16 PDT \\(~8 hours 18 minutes\\), with intermittent widget failures; system was restarted and under active monitoring from 08:25 PDT onward.\n\n### What was not impacted?\n\n* Data ingestion and processing\n* SEI 1.0 customers\n* Integrations and metadata flows\n\nNo customer data was lost.\n\n## Remediation\n\nUpon identifying the issue, the affected database nodes were restarted at 08:25 PDT, which restored initial stability. We continued to monitor the system closely, and when intermittent degradation was still observed afterward, we applied several additional fixes:\n\n* Adjusted database configuration settings to limit the amount of memory used for processing accumulated metadata, and tuned query-planning settings to reduce memory pressure.\n* Ran cleanup jobs to reduce the backlog of accumulated metadata on the affected tables.\n* Increased capacity on the affected database nodes to provide additional headroom.\n\nThese changes progressively stabilized the system, and the incident was fully resolved at 16:16 PDT.\n\n## Action Items\n\nTo prevent recurrence, we are implementing the following:\n\n1. We have upgraded the backend which includes underlying improvements that handle memory spikes caused by excessive delete files.\n2. We have rolled out automated  compaction jobs for newly introduced tables to prevent delete file accumulation going forward.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-22T16:16:51.506-07:00",
"resolved_inferred": false,
"started_at": "2026-07-22T07:55:15.285-07:00",
"state": "postmortem",
"title": "AIDI Dashboards \u2013 Degraded Performance",
"updated_at": "2026-08-10T11:24:01.842-07:00",
"url": "https://stspg.io/kzwzjz0p0wcp"
},
{
"body": "# Summary\n\nOn 17 July 2026, following a routine code deployment, customers on older delegate versions \\(858xx and below\\)  began experiencing delayed CI builds  on Harness Cloud-hosted builds using our global build-queueing capability. Affected builds experienced an unexpected pause of up to approximately 8 minutes at the \"waiting for infrastructure\" stage before continuing, rather than proceeding within the expected sub-second time. Overall build slowness was intermittent.\n\n# Impact\n\n* All CI builds were potentially subject to delay; impact was most pronounced for builds on Harness Cloud-hosted infrastructure using the global build-queueing feature.\n* Affected builds experienced an unexplained pause of up to approximately 8 minutes before continuing, followed by a slower \"cold start\" since a pre-reserved compute slot was not available \u2014 this presented to users as slow builds rather than build failures.\n* Accounts running on newer delegate versions \\(858xx and above\\) were not impacted.\n* No builds failed outright as a direct result of this issue, and no data was lost.\n\n# Root Cause\n\nThe root cause was an internal code change that inadvertently broke how a specific build-queueing record was read back from our database once builds that had already been queued under the previous version of the code encountered the newly deployed version. We resolved the immediate impact by cleaning up the affected records and reverting the underlying code change, and we are implementing several safeguards to prevent this class of issue from recurring.\n\n\u200c\n\n# Next Steps\n\nWe assess the risk of a similar recurrence as low as the following actions are being understaken.The specific code path that caused this incident has already been reverted, and we are implementing structural safeguards so that this general class of issue cannot recur, regardless of where in the codebase it might otherwise occur.\n\n| **Corrective / Preventive Action** |\n| --- |\n| Add explicit, stable identifiers to all internal data classes that get stored in our database, so that future internal code reorganizations cannot break the system's ability to read back previously stored records. |\n| Introduce rollback and backward-compatibility testing in our pre-production environment, specifically designed to catch this class of issue before it reaches production. |",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-17T14:56:27.511-07:00",
"resolved_inferred": false,
"started_at": "2026-07-17T10:16:32.070-07:00",
"state": "postmortem",
"title": "Degraded CI performance",
"updated_at": "2026-08-03T19:53:30.200-07:00",
"url": "https://stspg.io/yzlj2tnzgywd"
},
{
"body": "# Summary\n\nBetween June 19 and July 17, 2026, the built-in Git Clone step \u2014 and any pipeline step using the drone-git clone plugin \u2014 failed on ARM64 Kubernetes build infrastructure with the error exec /usr/local/bin/clone: exec format error. AMD64 \\(Intel/AMD\\) builds, Windows builds, and the VM containerless binary path were not affected.\n\nThe root cause was a defect in our internal image publishing process that caused ARM64-tagged drone-git images to actually contain AMD64 binaries. We identified and mitigated the issue the same day it was reported by reverting the drone-git image to the last known-good version. No customer action or configuration change was required.\n\n# Root Cause\n\nOn June 19, 2026, a security remediation restructured how the drone-git image is built. AMD64 builds were updated correctly, but the ARM64 build pipeline didn't build ARM64 files directly \u2014 it adapted the AMD64 build file via text substitution and compiled it on ARM64 infrastructure. The June 19 change altered the AMD64 file so that substitution silently no-op'd instead of failing, so the pipeline published an image tagged ARM64 whose Git Clone and Git LFS binaries were still compiled for AMD64.\n\n# Impact\n\n* Affected: The built-in Git Clone step, and any pipeline step using the drone-git clone plugin, running on ARM64 Kubernetes build infrastructure, across all accounts, between July 6 and July 17, 2026.\n* Symptom: Builds failed at the Git Clone step with exec /usr/local/bin/clone: exec format error.\n* Not affected: AMD64 \\(Intel/AMD\\) Kubernetes and VM builds, Windows builds, the VM containerless execution path, and our hardened image variant.\n\n# Mitigation\n\nWe reverted the drone-git image version used across all affected services to the last known-good release. This fully resolved the ARM64 execution failures; no customer configuration changes were required.\n\n# Next Steps\n\nTo prevent such issues from happening again.\n\n* Rebuild the ARM64 image publishing pipeline to build our dedicated ARM64 build files directly, rather than adapting the AMD64 build files.\n* Enhance automated post-publish validation to every image release: verify binary architecture matches the image tag, and run a functional smoke test before an image is considered releasable.\n* Expand automated test coverage to include ARM64 Kubernetes build scenarios.\n* Remove the affected intermediate image versions from circulation once the corrected release is validated.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-17T07:16:41.003-07:00",
"resolved_inferred": false,
"started_at": "2026-07-17T03:30:52.000-07:00",
"state": "postmortem",
"title": "CI git clone calls are failing for specific arm builds",
"updated_at": "2026-07-22T10:26:09.104-07:00",
"url": "https://stspg.io/pljx3l92jxgx"
},
{
"body": "# Summary\n\nBetween June 19 and July 17, 2026, the built-in Git Clone step \u2014 and any pipeline step using the drone-git clone plugin \u2014 failed on ARM64 Kubernetes build infrastructure with the error exec /usr/local/bin/clone: exec format error. AMD64 \\(Intel/AMD\\) builds, Windows builds, and the VM containerless binary path were not affected.\n\nThe root cause was a defect in our internal image publishing process that caused ARM64-tagged drone-git images to actually contain AMD64 binaries. We identified and mitigated the issue the same day it was reported by reverting the drone-git image to the last known-good version. No customer action or configuration change was required.\n\n# Root Cause\n\nOn June 19, 2026, a security remediation restructured how the drone-git image is built. AMD64 builds were updated correctly, but the ARM64 build pipeline didn't build ARM64 files directly \u2014 it adapted the AMD64 build file via text substitution and compiled it on ARM64 infrastructure. The June 19 change altered the AMD64 file so that substitution silently no-op'd instead of failing, so the pipeline published an image tagged ARM64 whose Git Clone and Git LFS binaries were still compiled for AMD64. \n\n# Impact\n\n* Affected: The built-in Git Clone step, and any pipeline step using the drone-git clone plugin, running on ARM64 Kubernetes build infrastructure, across all accounts, between July 6 and July 17, 2026.\n* Symptom: Builds failed at the Git Clone step with exec /usr/local/bin/clone: exec format error.\n* Not affected: AMD64 \\(Intel/AMD\\) Kubernetes and VM builds, Windows builds, the VM containerless execution path, and our hardened image variant.\n\n# Mitigation\n\nWe reverted the drone-git image version used across all affected services to the last known-good release. This fully resolved the ARM64 execution failures; no customer configuration changes were required.\n\n# Next Steps\n\nTo prevent such issues from happening again.\n\n* Rebuild the ARM64 image publishing pipeline to build our dedicated ARM64 build files directly, rather than adapting the AMD64 build files.\n* Enhance automated post-publish validation to every image release: verify binary architecture matches the image tag, and run a functional smoke test before an image is considered releasable.\n* Expand automated test coverage to include ARM64 Kubernetes build scenarios.\n* Remove the affected intermediate image versions from circulation once the corrected release is validated.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-16T00:01:23.936-07:00",
"resolved_inferred": false,
"started_at": "2026-07-15T23:40:19.000-07:00",
"state": "postmortem",
"title": "Code services are degraded, git clone is failing",
"updated_at": "2026-07-22T10:13:19.239-07:00",
"url": "https://stspg.io/qgw0wh6qsxr4"
},
{
"body": "## Summary\n\nOn July 14, 2026, certain Harness Feature Flag \\(FF\\) Classic customers on the Prod1 and Prod2 production environments were unable to view Feature Flags in the Harness platform. Affected requests returned an authorization error \\(HTTP 403\\), so Feature Flags were not visible for those users until access was restored.\n\nThe behaviour was caused by a planned security update that began requiring an additional Feature Flags permission for related read operations. Users whose roles did not yet include that permission were correctly denied access, which appeared as a product failure. Harness temporarily rolled back the Feature Flags service change to restore access, updated the required permissions for affected users and roles, and confirmed that Feature Flag visibility returned to normal. The stronger permission checks remain in place going forward.\n\n## Impact\n\nDuring the customer-reported incident window on July 14, 2026 \\(status page approximately 11:25 UTC to 11:57 UTC\\):\n\n* Certain Feature Flag Classic customers on **Prod1** and **Prod2** were impacted.\n* Affected users could not view Feature Flags in the Harness UI / API and received **403 Forbidden** responses.\n* Impact was limited to users and roles that did not yet have the updated Feature Flags permission required by the security change.\n* A public status page update was posted for Prod1 and Prod2 Feature Flags.\n\nThere was **no data loss**, no change to stored feature flag configurations, and no impact to Feature Flag evaluation for applications whose SDK credentials and permissions were unaffected. Customers and users with the required permission continued to operate normally.\n\n## Root Cause\n\nHarness deployed a planned Feature Flags authorization update that enforces the `ff_targetgroup_view` permission on target-related read paths used when viewing Feature Flags. This enforcement is intentional and remains the expected behaviour.\n\nUsers and roles that had not yet been granted `ff_targetgroup_view` received 403 responses and could not see Feature Flags. From the customer\u2019s perspective this looked like an outage; it was an authorization denial due to missing required permissions after the security update.\n\n## Mitigation\n\nHarness completed the following mitigation steps:\n\n* Temporarily rolled back the Feature Flags service in Prod2 and Prod1 \\(and aligned Prod0\\) to restore access quickly while permissions were corrected.\n* Updated / granted the required `ff_targetgroup_view` permission for affected users and roles.\n* Verified that Feature Flag visibility returned to normal and closed the status page incident.\n\nThese actions restored customer access. The permission enforcement introduced by the security update **remains in effect** going forward; the lasting fix is correct permission assignment, not removal of the check.\n\n## Preventive Actions\n\nTo avoid a recurrence of this issue, Harness is taking the following actions:\n\n* Keeping the stronger Feature Flags permission checks in place as the permanent security posture.\n* Ensuring role and permission updates for `ff_targetgroup_view` are applied with \\(or before\\) similar authorization changes in production.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-14T04:57:39.917-07:00",
"resolved_inferred": false,
"started_at": "2026-07-14T04:25:41.916-07:00",
"state": "postmortem",
"title": "Certain users are unable to see feature flags in prod1 and prod2",
"updated_at": "2026-07-23T15:32:36.677-07:00",
"url": "https://stspg.io/68m3d8z9lnsk"
},
{
"body": "# Summary\n\nOn July 9, 2026, some customers in the Harness Prod3 cluster experienced pipeline failures over a ~2.5-hour window, despite making no changes in their pipelines. Shell Script steps failed when fetching scripts from the Harness File Store, and Kubernetes deployment steps failed when retrieving secret encryption details.\n\n# Impact\n\n1. **Affected users**: Some customers using the Harness Prod3 cluster reported pipeline execution failures.\n2. **User impact:** Customers experienced failures in Shell Script and Kubernetes deployment steps. Impacted pipelines required manual retries after the incident was resolved.\n3. **Scope:** This was not a platform-wide outage. The issue was isolated to specific APIs in our internal service within the Prod3 cluster.\n\n# Root Cause\n\nThe incident was caused by degraded performance in a downstream service. This led to a storm of timeout errors in our internal service, which caused these failures.\n\n# Mitigation\n\nEngineering identified the change that caused this issue in the downstream dependency service and deployed a patch to restore it. Once deployed, pipeline executions returned to normal success rates.\n\n# Preventive Measures & Next Steps\n\n1. Added additional monitoring and alerting on the downstream service to detect performance degradation earlier.\n2. Reviewing timeout and retry logic in the internal service to improve resilience against downstream latency spikes.\n3. Evaluating circuit breaker patterns to prevent cascading failures from downstream dependencies.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-09T14:46:38.786-07:00",
"resolved_inferred": false,
"started_at": "2026-07-09T13:21:53.383-07:00",
"state": "postmortem",
"title": "Few customers experiencing intermittent errors during pipeline execution.",
"updated_at": "2026-07-20T22:21:21.415-07:00",
"url": "https://stspg.io/fv5drkv33y06"
},
{
"body": "## Summary\n\nBetween July 7 and July 8, 2026, customer accounts hosted in our Prod 3 environment experienced intermittent slowness and failures when loading pipelines and resolving templates. The underlying cause was a sharp, sustained increase in internal feature-flag lookup traffic from our template-processing service to our core platform backend service.\n\nThe issue presented across two separate days. On Day 1 \\(July 7\\), our team restored service through infrastructure level mitigations by restarting affected services and adding capacity. Because that underlying cause was still present, the same failure mode recurred on Day 2 \\(July 8\\). Deeper investigation on Day 2 identified the specific feature flag and code path responsible, and engineering shipped a fix so that flag check is now served from local cache. This fix was deployed across all production clusters and fully resolved the issue; it has not recurred since\n\nWe are auditing all other feature flags on the platform for the same caching gap, and are implementing process and monitoring improvements described in the Preventive Measures section below to catch this class of issue earlier and reduce the risk of recurrence.\n\n## Impact\n\n**Symptoms:** Intermittent slowness or failure loading pipelines; slow or failed template validation and resolution; intermittent connectivity/timeout errors between platform services.\n\n**Functional impact:** Pipeline executions and pipeline-studio operations that depend on template resolution were delayed or failed during the active windows described in the timeline above. Some executions required manual retry\n\n## Root cause\n\nOur core backend service could not keep up with the request arrival rate because of a sustained, uncached feature-flag lookup driving a large and growing volume of internal traffic following a recent release.\n\n\u200c\n\n## Mitigation\n\n* Immediate mitigation \\(Day 1 and Day 2\\): removed the saturated backend instance\\(s\\) from service and added capacity to relieve the acute bottleneck.\n* Permanent fix \\(Day 2\\): updated the template-processing service so the responsible feature-flag check reads from local cache instead of calling the backend service on every lookup, eliminating the excess internal traffic at its source.\n* The fix was deployed to all production clusters and verified through a return to normal internal traffic volumes and response times.\n\n## Preventive Measures & Next Steps\n\n* Feature-flag caching audit: We are reviewing all feature flags used in high-frequency code paths across the platform to identify and close any similar caching gaps before they can cause a repeat of this issue.\n* Release validation: We are strengthening pre-release load testing for new feature-flag-gated code paths so that call-volume regressions of this kind are caught before reaching production.\n* Capacity alerting: We are adding proactive alerting on backend request-queue saturation so that this class of bottleneck is caught and mitigated automatically, before it affects customer-facing response times.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-08T11:56:41.291-07:00",
"resolved_inferred": false,
"started_at": "2026-07-08T00:49:30.000-07:00",
"state": "postmortem",
"title": "Prod3: Slowness in Harness Platform",
"updated_at": "2026-07-13T17:14:31.957-07:00",
"url": "https://stspg.io/gvfsbhwjk485"
},
{
"body": "## Summary\n\nBetween July 7 and July 8, 2026, customer accounts hosted in our Prod 3 environment experienced intermittent slowness and failures when loading pipelines and resolving templates. The underlying cause was a sharp, sustained increase in internal feature-flag lookup traffic from our template-processing service to our core platform backend service.\n\nThe issue presented across two separate days. On Day 1 \\(July 7\\), our team restored service through infrastructure level mitigations  by  restarting affected services and adding capacity. Because that underlying cause was still present, the same failure mode recurred on Day 2 \\(July 8\\). Deeper investigation on Day 2 identified the specific feature flag and code path responsible, and engineering shipped a fix so that flag check is now served from local cache. This fix was deployed across all production clusters and fully resolved the issue; it has not recurred since\n\nWe are auditing all other feature flags on the platform for the same caching gap, and are implementing process and monitoring improvements described in the Preventive Measures section below to catch this class of issue earlier and reduce the risk of recurrence.\n\n## Impact\n\n**Symptoms:** Intermittent slowness or failure loading pipelines; slow or failed template validation and resolution; intermittent connectivity/timeout errors between platform services.\n\n**Functional impact:** Pipeline executions and pipeline-studio operations that depend on template resolution were delayed or failed during the active windows described in the timeline above. Some executions required manual retry\n\n## Root cause\n\nOur core backend service could not keep up with the request arrival rate because of a sustained, uncached feature-flag lookup driving a large and growing volume of internal traffic following a recent release.\n\n\u200c\n\n## Mitigation\n\n* Immediate mitigation \\(Day 1 and Day 2\\): removed the saturated backend instance\\(s\\) from service and added capacity to relieve the acute bottleneck.\n* Permanent fix \\(Day 2\\): updated the template-processing service so the responsible feature-flag check reads from local cache instead of calling the backend service on every lookup, eliminating the excess internal traffic at its source.\n* The fix was deployed to all production clusters and verified through a return to normal internal traffic volumes and response times.\n\n## Preventive Measures & Next Steps\n\n* Feature-flag caching audit: We are reviewing all feature flags used in high-frequency code paths across the platform to identify and close any similar caching gaps before they can cause a repeat of this issue.\n* Release validation: We are strengthening pre-release load testing for new feature-flag-gated code paths so that call-volume regressions of this kind are caught before reaching production.\n* Capacity alerting: We are adding proactive alerting on backend request-queue saturation so that this class of bottleneck is caught and mitigated automatically, before it affects customer-facing response times.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-07T05:36:37.965-07:00",
"resolved_inferred": false,
"started_at": "2026-07-07T03:22:06.839-07:00",
"state": "postmortem",
"title": "Prod3: Slowness in loading pipeline components",
"updated_at": "2026-07-13T17:14:14.140-07:00",
"url": "https://stspg.io/0w316hfsy8n6"
},
{
"body": "**Incident Summary**\n\nHarness FME \\(Feature Management & Experimentation\\) faced substantial delays in scheduled impressions data exports on Jul 2, 2026. This issue stemmed from performance degradation within Tinybird's infrastructure, affecting query execution crucial to our export pipeline.\n\n**Root Cause**\n\nThe core issue was a temporary **performance degradation on the provider's infrastructure**. This directly impacted the metadata retrieval endpoint, which failed to return job IDs intermittently. Consequently, completed export jobs underwent repeated retries, leading to a backlog in the export queue.\n\n**Impact**\n\n* The export backlog accumulated, peaking at more than 6 hours behind the planned schedule.\n* No data loss was detected during the incident.\n\n\u200c\n\n**Mitigation**\n\n1. **Adjusted Temporal workflows:** Extended timeouts from ~30 to ~50 minutes and increased polling attempts from ~180 to ~300.\n2. **Expanded resource capacity.**\n3. **Traffic redirection:** Systematically shifted export workloads over to an optimized in-house cluster.\n4. **Manual intervention:** Terminated stalling workflows stuck in retry loops from failed metadata lookups.\n\n\u200c\n\n**Next Steps**\n\nComplete the migration of all remaining export jobs to the internal analytics cluster \\(In progress\\).",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-03T09:11:22.765-07:00",
"resolved_inferred": false,
"started_at": "2026-07-02T08:48:25.564-07:00",
"state": "postmortem",
"title": "FME - Some customers are experiencing delays in scheduled exports of impressions",
"updated_at": "2026-07-09T12:34:44.544-07:00",
"url": "https://stspg.io/pydwzxnfsc2x"
},
{
"body": "### **Summary**\n\nDuring a planned security improvement to rotate service authentication secrets in the Prod 4 environment, an incorrect secret value was inadvertently applied to a subset of internal services. This resulted in authentication failures between platform components, impacting CI initialization, delegate task creation, and log streaming for certain pipeline executions. The issue was resolved by rolling back the token configuration to the previous known-good version and validating service recovery.\n\n### **Impact**\n\n**Affected environment:** Prod 4\n\n**Customer-visible impact:**\n\n* CI builds failed during initialization with errors such as:\n\n    * _\u201cCould not fetch token from log service\u201d_\n    * _\u201ctoken in request not authorized for receiving tokens\u201d_ \\(HTTP 400\\)\n    \n* CD and IaCM pipeline stages continued to execute but did not stream logs.\n* Some Shell Script steps failed during execution.\n* Delegate tasks could not be created or assigned for some pipeline executions.\n\nNo customer data was lost or corrupted.\n\n### **Root Cause**\n\nAs part of a planned security initiative, authentication secrets used for communication between internal platform services were rotated in the Prod 4 environment.\n\nDuring the rotation process, an incorrect secret value was inadvertently configured for one of the services. Instead of applying the newly generated secret in the expected format, a default/incorrect value was introduced. This created a mismatch between services that authenticate using the shared token, causing authentication requests to be rejected with HTTP 400 errors.\n\nThe authentication failures prevented CI components from obtaining log service tokens, disrupted delegate task creation, and prevented log streaming for affected pipeline executions.\n\n### **Resolution**\n\nEngineering responded by:\n\n* Reverting the token configuration to the previous known-good secret.\n* Restoring consistent authentication across affected services.\n* Validating recovery through CI pipeline execution tests.\n* Verifying CD and IaCM pipeline execution and log streaming.\n* Executing a comprehensive validation suite to confirm platform functionality before closing the incident.\n\n### **Follow-up Actions**\n\nTo reduce the likelihood of similar incidents in the future, Harness will:\n\n* Implement additional pre-rotation and post-rotation validation checks to verify that the correct secret values and formats are applied before changes are activated.\n* Introduce automated verification of authentication between dependent services immediately following secret rotations.\n* Strengthen operational safeguards and deployment validation for security credential rotation procedures to detect configuration mismatches before they can impact customer workloads.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-06-29T08:41:18.884-07:00",
"resolved_inferred": false,
"started_at": "2026-06-29T07:56:24.965-07:00",
"state": "postmortem",
"title": "Few pipelines in prod4 are running with degraded performance",
"updated_at": "2026-07-01T17:25:57.690-07:00",
"url": "https://stspg.io/0cs7fd3lmd1n"
},
{
"body": "## **Summary**\n\n\u200c\n\n_Between **06:50 UTC and 07:20 UTC on June 26, 2026**, customers experienced intermittent login failures while accessing Traceable environments across multiple US clusters. The incident was traced to an issue with the external authentication provider \\(Auth0\\), where elevated socket timeouts caused increased login latency and authentication request failures. Login functionality gradually recovered as the upstream issue stabilized, and the incident was resolved after successful login validation across multiple impacted clusters._\n\n## **Root Cause**\n\n_The root cause was an issue with the external authentication provider \\(Auth0\\), which experienced elevated socket timeouts while processing authentication requests. These upstream timeouts increased login latency and caused intermittent authentication failures across multiple US-region clusters. Since authentication requests depended on the external provider, affected login attempts failed despite Traceable platform services remaining healthy. The incident was resolved once the upstream authentication service recovered and login requests consistently completed successfully._\n\n## **Impact**\n\n_Starting at approximately 06:50 UTC, customers experienced intermittent login failures when accessing Traceable environments across multiple US clusters. The issue affected user authentication, preventing some users from accessing the platform while underlying application services remained operational. Login functionality progressively recovered during the incident, and normal authentication was restored by 07:20 UTC._\n\n## **Remediation**\n\n_The engineering team worked with the external authentication provider while continuously monitoring authentication health across affected clusters. Login functionality was validated through platform metrics and manual verification across representative environments. After confirming consistent authentication success across impacted clusters, the incident was declared resolved._\n\n## **Action Items**\n\n_To prevent such issues going forward, Harness will,_\u00a0\u00a0\n\n_Increase authentication resilience: Evaluate improvements to authentication request handling, including timeout tuning, retry strategies where appropriate, and graceful degradation for transient upstream failures._",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-06-26T00:31:56.238-07:00",
"resolved_inferred": false,
"started_at": "2026-06-26T00:20:39.157-07:00",
"state": "postmortem",
"title": "Login to Traceable Clusters Impacted",
"updated_at": "2026-06-30T10:13:51.261-07:00",
"url": "https://stspg.io/dwdctr7c33dt"
},
{
"body": "### **Summary**\n\nDuring a planned security improvement to rotate service authentication secrets in the EU1 environment, an out of sequence  secret value was inadvertently applied to a subset of internal services. This resulted in authentication failures between platform components, causing login failures. The issue was resolved by rolling back the token configuration to the previous known-good version and validating service recovery.\n\n### **Impact**\n\n**Affected environment:** EU1\n\n**Customer-visible impact: Login Failures** \n\nNo customer data was lost or corrupted.\n\n### **Root Cause**\n\nAs part of a planned security initiative, authentication secrets used for communication between internal platform services were rotated in the EU1 .\n\nDuring the rotation process, an  secret value was inadvertently misconfigured  for one of the services. There was a newer format of secret and had to be applied in a certain order.  This created a mismatch between services that authenticate using the shared token, causing authentication requests to be rejected with HTTP 400 errors.\n\nThe authentication failures prevented CI components from obtaining log service tokens, disrupted delegate task creation, and prevented log streaming for affected pipeline executions.\n\n### **Resolution**\n\nEngineering responded by:\n\n* Reverting the token configuration to the previous known-good secret.\n* Restoring consistent authentication across affected services.\n\n### **Follow-up Actions**\n\nTo reduce the likelihood of similar incidents in the future, Harness will:\n\n* Implement additional pre-rotation and post-rotation validation checks to verify that the correct secret values and formats are applied before changes are activated.\n* Introduce automated verification of authentication between dependent services immediately following secret rotations.\n* Strengthen operational safeguards and deployment validation for security credential rotation procedures to detect configuration mismatches before they can impact customer workloads.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-06-24T07:37:42.747-07:00",
"resolved_inferred": false,
"started_at": "2026-06-24T05:03:14.000-07:00",
"state": "postmortem",
"title": "Harness Login Failure in EU1 environment",
"updated_at": "2026-07-01T17:41:27.244-07:00",
"url": "https://stspg.io/cf1x39hmlr8w"
},
{
"body": "# **Customer RCA: Executive Summary**\n\nOn June 17, 2026, between 9:30 PM and 10:15 PM IST, customers in certain production environments experienced failures when accessing IDP \\(Internal Developer Portal\\) workflows. The issue was caused by a configuration mismatch in service authentication settings.\n\n## **Impact**\n\n* **Affected Service:** IDP workflows\n* **Duration:** Approximately 45 minutes \\(9:30 PM - 10:15 PM IST\\)\n* **Customer Impact:** Users were unable to load IDP workflows during this period.\n* **Data Loss:** None\n* **Security Impact:** None\n\n## **Root Cause**\n\nDuring a planned service deployment, a configuration mismatch occurred in the authentication settings used for internal service-to-service communication. Specifically, the authentication credentials configured in the newly deployed service did not match the credentials expected by existing services in prod environments.\n\nThis mismatch caused authentication failures when services attempted to communicate with each other, resulting in workflow loading failures for customers.\n\n## **Mitigation**\n\nThe issue was resolved by reverting the configuration change to restore the previous working authentication settings. The engineering team was proactively monitoring logs immediately after deployment and identified error patterns several minutes before the first customer report.\n\n## **Action Items/Next Steps**\n\n1. **Enhanced Pre-Deployment Validation:** Automated validation checks will be introduced to verify that all required configurations, secrets, and service dependencies are correctly provisioned and consistent across services before any production deployment.\n2. **Improved Deployment Procedures:** Updated deployment checklists with explicit verification of service authentication settings across services before rollout.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-06-17T10:50:54.278-07:00",
"resolved_inferred": false,
"started_at": "2026-06-17T09:44:50.009-07:00",
"state": "postmortem",
"title": "IDP Workflows and Backend failures",
"updated_at": "2026-06-23T17:53:26.755-07:00",
"url": "https://stspg.io/3khs7kvsfycw"
},
{
"body": "## Summary\n\nOn June 17, 2026, a subset of Harness Hosted CI customers experienced intermittent build failures when downloading Java/Gradle dependencies from Maven Central. Affected builds failed with `429 Too Many Requests` errors and could not proceed until the issue was resolved. \n\nThe incident was not caused by a Harness code change or platform defect. It was triggered by a policy change implemented by Sonatype \\(the operator of Maven Central\\) that tightened IP-based rate limits for unauthenticated access. Because Harness Hosted CI routes builds through shared network egress infrastructure, the aggregate outbound traffic from multiple customers sharing a common IP address exceeded the new, lower thresholds \u2014 causing all builds behind that IP to be temporarily blocked.\n\nThis was an industry-wide change. Other CI/CD platforms experienced identical incidents in the weeks prior to Harness being affected.\n\n## Root Cause\n\n**Primary cause:** Sonatype tightened IP-based rate limiting policies on Maven Central \\([repo1.maven.org](http://repo1.maven.org) and [repo.maven.apache.org](http://repo.maven.apache.org)\\). Harness Hosted CI uses shared network egress infrastructure where multiple customer build environments route outbound traffic through common IP addresses. When the aggregate Maven Central request volume from all customers behind a given IP exceeded Sonatype's revised threshold, that IP was hard-blocked for approximately 30 minutes, causing all builds attempting to download Maven dependencies to fail with `429` errors.\n\n\u200c\n\n## Mitigation\n\n**Immediate actions taken during the incident:**\n\n1. **Regional redistribution:** For customers experiencing active failures, builds were routed to infrastructure regions whose outbound IPs had not yet hit the rate limit threshold, providing immediate relief for those customers.\n2. **Sonatype engagement and IP whitelisting:** The Harness team contacted Sonatype support and provided the full list of Harness Hosted CI outbound IP addresses for whitelisting. Sonatype applied a \"warning\" rate-limit policy in place of hard blockingeliminating the `429` errors across all affected regions. Post-mitigation validation confirmed zero throttling across all tested IPs.\n\n## Next Steps\n\nThe following actions are in progress or planned to prevent recurrence and reduce exposure to similar incidents:\n\n**Near-term \\(in progress\\):**\n\n*  Harness is in the process of establishing a commercial agreement with Sonatype for elevated rate limits on Maven Central and advance notification of future policy changes.\n* Harness is engaging other major dependency registries \u2014 including Docker Hub, npm, GitHub Packages, and the Gradle Plugin Portal \u2014 to proactively whitelist Harness egress IPs before similar incidents can occur.\n* We are publishing documentation on best practices for managing external dependency access in Hosted CI, including guidance on dependency proxies, mirrors, and caching strategies.\n*  Harness plans to implement caching infrastructure for common dependencies at the platform level \\(using Harness Artifact Registry\\), which will significantly reduce direct Maven Central request volume and provide resilience against future external rate-limit changes.\n* We are enhancing our alerting alerting for `429` error patterns from external registries so that future rate-limit incidents are detected internally before customers are impacted.\n\n\u200c\n\n## What This Means for You\n\n**To reduce future exposure**, we recommend:\n\n* **Enable Cache Intelligence** in Harness CI to cache dependency artifacts between builds, reducing how often Maven Central is contacted.\n* **Use Harness Save/Restore Cache steps** to persist your local dependency cache across pipeline runs.\n\nIf you have questions or need help implementing any of these recommendations, please contact your Customer Success Manager or reach out to Harness Support.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-06-17T10:13:14.903-07:00",
"resolved_inferred": false,
"started_at": "2026-06-17T06:42:40.254-07:00",
"state": "postmortem",
"title": "Cloud builds are failing with 429 Too Many Requests error while downloading artifacts from central sonatype registry.",
"updated_at": "2026-06-23T16:56:08.320-07:00",
"url": "https://stspg.io/m04qndlxz5fm"
},
{
"body": "## Summary\n\nOn June 8, 2026, a subset of Harness customers experienced UI degradation in the Harness Code Repository module. Affected users saw broken styling \u2014 most notably, the Create Repository control was visually impaired and difficult to use. The issue was not a full outage; core functionality remained available\n\nUsers who logged out and back in, or opened the application in an incognito/private window, were not affected \u2014 pointing to browser-stored state as the failure mechanism. \n\n## Root Cause\n\n**Primary cause:** A change to the Harness Platform UI framework altered the default theme fallback from a valid, deployable CSS theme to a \"system preference\" token \u2014 a value that instructs the browser to follow the operating system's light/dark mode setting. While this token is valid as a user _preference_, it is not itself a CSS theme: no corresponding stylesheet bundle exists for it. When users without a previously saved theme preference loaded the Harness UI, this unresolvable token was immediately written to their browser's local storage and applied to the page. The result was missing CSS design tokens, which broke the visual rendering of Harness Code.\n\nBecause the invalid value was persisted in browser storage, it continued affecting those users across subsequent page navigations until their session was cleared.\n\n\u200c\n\n## Mitigation\n\n**Immediate actions:**\n\n The Harness team rolled back the Platform UI to the last known good version on all affected production environments. This stopped new users from hitting the broken default and, once users cleared their session \\(via logout/login\\), restored correct rendering.\n\nAffected users were advised to log out and log back in to clear the stale value from browser storage. Opening a new incognito/private window also restored correct behavior immediately.\n\nA targeted fix reverting the default theme to the correct value was merged and deployed. This permanently addressed the root cause without requiring users to take any action.\n\n## Next Steps\n\nThe following actions are in progress or planned to prevent recurrence:\n\n\u200c\n\n* **Default theme reverted:** The platform UI wrapper now uses a valid, CSS-resolvable theme as the default fallback. `system` preference tokens will no longer be used as raw defaults. \\(Done\\)\n* **Browser storage cleanup:** A cleanup was deployed to migrate any stale invalid theme values already persisted in affected users' browsers, resolving the issue without requiring manual logout. \\(Done\\)\n* **Defensive validation in the theme provider:** Adding input validation so that any theme value \u2014 whether loaded from storage or applied as a default \u2014 is checked against the set of known valid CSS themes before being written to storage or applied to the DOM. Invalid values will fall back safely to a known-good theme.\n* **Expanded test coverage:** Adding automated end-to-end tests that exercise the theme initialization path for users across both the new and legacy UI experiences, including first-time users with no previously saved preference.\n* **Client-side monitoring for theme resolution failures:** Adding telemetry to detect when an invalid theme value is applied, enabling the team to catch styling regressions before they are reported by customers.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-06-08T12:50:59.538-07:00",
"resolved_inferred": false,
"started_at": "2026-06-08T10:32:27.651-07:00",
"state": "postmortem",
"title": "Certain customers may see odd web artifacts on Code UI. Logging out and Logging in resolves this issue.",
"updated_at": "2026-06-23T17:42:33.100-07:00",
"url": "https://stspg.io/ldcxndjc9ryy"
},
{
"body": "## **Summary**\n\n_Between 13:30 UTC on June 5, 2026, and 00:27:00 UTC on June 6, 2026, customers using Harness IDP experienced intermittent 503 errors when accessing IDP web pages and APIs. The incident was mitigated through a configuration update and normal operation was restored by 00:27 UTC._\n\n## **Root Cause**\n\n_The Root Cause was traced back to a performance optimization update deployed around 10:35 UTC on June 5, 2026. This network transport configuration change altered idle TCP connection behavior, causing long-lived connections to be terminated prematurely under certain conditions. Consequently, requests failed to reach backend services, leading to intermittent 503 errors in the Harness IDP service. Stability was restored by recalibrating connection timeouts and lifecycle settings to synchronize connection management across the network path and eliminate stale connection reuse._\n\n## **Impact**\n\n_Starting at approximately ~13:30 UTC, customers experienced intermittent 503 service errors while using the Harness Internal Developer Portal \\(IDP\\) until we resolved the issue at 00:27:00 UTC. The issue impacted both UI and API traffic, resulting in failed requests and a degradation of service availability until mitigation measures were implemented._\n\n## **Remediation**\n\n_The team implemented a configuration change to connection management settings, reducing idle connection timeouts and limiting connection lifetime._\u00a0\n\n_This prevented the reuse of stale connections and restored service stability._\n\n## **Action Items**\n\n1. _**Enhance monitoring & alerting** \u2013 Enhance synthetic alerting and coverage for upstream connection failures, connection resets, and elevated 503 error rates to enable earlier detection of network transport issues._\n2. _**Pre-production validation of networking changes** \u2013 Expand pre-production checks for configuration alignment and testing for network timeout and connection management scenarios to reduce the risk of similar issues in the future._",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-06-05T17:30:38.941-07:00",
"resolved_inferred": false,
"started_at": "2026-06-05T15:56:19.954-07:00",
"state": "postmortem",
"title": "Intermittent 503s on IDP services",
"updated_at": "2026-06-19T03:52:26.803-07:00",
"url": "https://stspg.io/clncyx4rm5x9"
},
{
"body": "### Summary\n\nOn June 4, 2026, customers using Harness FME SDKs experienced degraded streaming performance. SDK connections repeatedly disconnected and fell back to polling mode, resulting in delayed flag propagation and elevated log noise. Flag evaluation itself was unaffected \u2014 all SDKs are designed to poll against the CDN during streaming disconnects. \n\n\u200c\n\n\u200c\n\n### Impact\n\n* **Region affected:** us-east-1 \\(other regions \u2014 EU, APAC \u2014 remained healthy\\)\n* **Customer impact:** SDK streaming connections intermittently dropped, causing SDKs to fall back to polling. Flag evaluation continued uninterrupted; customers experienced delayed flag propagation\n\n### Root Cause\n\nThe underlying cause was memory exhaustion in the backend messaging cluster in the us-east-1 region, triggered by a combination of an ongoing service rollout and a traffic spike during the same . Other regions remained fully healthy throughout. \n\n### Remediation Actions Taken\n\n1.  Service was rolled back to the previously stable version.\n2. Increased capacity to handle the surge of connections.\n\n### Preventive Measures / Next Steps\n\nTo prevent such issues happening again, Harness is working on several optimizations:\n\n* **Enhance** deployments to do validation before broader rollout to catch capacity pressure early.\n* **Improve** connection handling during deployments or more conservative deployment pacing during high-load periods. \n* **Improved SDK behavior on reconnect**:  Connection recycles inherently generate a resync storm; further investigation ongoing into reducing the cascading request impact from mass reconnects",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-06-04T09:38:04.620-07:00",
"resolved_inferred": false,
"started_at": "2026-06-04T07:33:42.868-07:00",
"state": "postmortem",
"title": "FME SDKs - SDKs falling back to polling",
"updated_at": "2026-07-01T16:32:51.370-07:00",
"url": "https://stspg.io/ld5lhpnjsrf3"
},
{
"body": "Similar to the incident prior to this on 05/27",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-06-01T06:47:22.225-07:00",
"resolved_inferred": false,
"started_at": "2026-06-01T03:01:49.734-07:00",
"state": "postmortem",
"title": "Hosted CI Builds Failing Intermittently",
"updated_at": "2026-06-24T15:41:42.920-07:00",
"url": "https://stspg.io/t7prjq3p6pv6"
},
{
"body": "## Summary\n\nOn May 27, 2026, customers using Harness Pipelines on the legacy log-storage backend experienced intermittent failures and missing pipeline logs in the UI. **No log data was permanently lost**\n\n## Root cause\n\nThe root cause was CPU saturation of the log service's underlying cache infrastructure, triggered by an automation script from one enterprise customer that opened a large number of long-lived log-streaming connections without closing them.\n\n## Mitigation\n\n**Immediate actions taken during the incident:**\n\n\u200c\n\n1. **Rate limiting applied** Traffic from the automation generating the runaway connections was rate-limited at the network layer, reducing new connection creation and allowing the CPU to begin recovering.\n2. **Streaming connection timeout reduced:** The maximum lifetime for legacy streaming connections was reduced from 1 hour to 10 minutes across production environments, limiting how long any single connection can remain open and reducing the steady-state connection count.\n3. **Affected customer migrated to newer storage infra:** The enterprise customer most affected by the incident was migrated to the new log storage path infrastructure. Once migrated, their logs were immediately accessible, confirming that no log data had been lost. The cache CPU returned to normal levels \\(~90% sustained CPU dropped\\) following this migration.\n\n## Next Steps\n\nThe following actions are in progress or planned to prevent recurrence:\n\n* **Rate limiting applied** for the specific traffic pattern that triggered this incident. \\(Done\\)\n* **Streaming connection timeout reduced to 10 minutes** across primary production environments. \\(Done\\)\n* **Complete migration to new log storage infra :** The primary long-term fix is finishing the migration of all remaining accounts \n\n* **Enhance monitoring alerts:** A broader audit of all monitoring alert rules is underway to identify any other critical detection paths that may be silently disabled.\n* **Extend streaming connection timeout reduction globally:** The 10-minute maximum streaming connection lifetime is being rolled out to all remaining environments \\(including development and free-tier clusters\\) to ensure consistent protection.\n* **Per-account rate limiting on streaming connections:** We are adding an application-level cap on concurrent streaming connections per account. The initial implementation will be observe-only \\(logging warnings when thresholds are approached\\) to gather data on real-world usage before converting to hard enforcement.\n* **Improved error messaging:** When log operations fail due to underlying infrastructure issues \\(cache timeouts, I/O errors\\), the error surfaces to users and logs will be updated to accurately reflect the infrastructure cause, rather than showing a generic \"stream not found\" message that obscures the true reason for failure.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-05-27T16:20:35.975-07:00",
"resolved_inferred": false,
"started_at": "2026-05-27T07:20:18.099-07:00",
"state": "postmortem",
"title": "Pipeline logging streams experiencing degraded delivery to console.",
"updated_at": "2026-06-23T20:08:06.838-07:00",
"url": "https://stspg.io/2xvs1403yxlh"
},
{
"body": "### **Summary**\n\nBetween May 27 and June 1, 2026, some Harness CI customers experienced pipeline execution failures with the error:\n\n`failed to call LE.RetryStartStep: context deadline exceeded`\n\n###  **Impact**\n\nAffected customers saw intermittent CI pipeline failures during step execution.\n\n Existing pipeline definitions, customer data, source code, and artifacts were not impacted.\n\n### **Root Cause**\n\nThe root cause was a deadlock in the Light Engine logging path. When the log service returned an error, the Light Engine log writer attempted to reacquire a mutex it already held. This caused the Light Engine process to freeze, which led to step execution timeouts and pipeline failures. \u00a0\n\n### **Mitigation and Resolution**\n\nHarness Engineering took multiple mitigation steps during the incident, including:\n\n* Rolled back affected runner versions where needed\n* Increased relevant timeout configurations\n* Reduced log-service load and latency\n* Temporarily disabled the affected livelog streaming path\n* Migrated selected workloads across regions and infrastructure providers\n* Pinned a fixed Light Engine version through runner configuration\n\nThe final fix addressed the mutex deadlock in the Light Engine log writer and prevented the same lock from being reacquired while already held. \u00a0\n\n### **Prevention and Follow-Up Actions**\n\nHarness is taking the following actions to reduce recurrence risk:\n\n* Improve deadlock detection in critical concurrent code paths\n* Strengthen error handling for log-service interactions\n* Add better monitoring for Light Engine process health\n* Improve safeguards around logging-path failures\n* Continue reviewing runner rollout and validation processes",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-05-27T18:48:45.368-07:00",
"resolved_inferred": false,
"started_at": "2026-05-27T05:35:32.385-07:00",
"state": "postmortem",
"title": "Hosted CI Builds Failing Intermittently",
"updated_at": "2026-06-04T08:47:06.927-07:00",
"url": "https://stspg.io/ysb2cz4hly1v"
},
{
"body": "# **Summary**\n\nOn May 20 and May 21, 2026, customers on the Prod3 cluster experienced degraded pipeline performance on two separate occasions. In both cases, pipeline executions ran slower than usual and execution graph rendering was delayed because database operations were being queued on the backing MongoDB instance. Both incidents were mitigated by adjusting database write settings and adding capacity headroom, which relieved pressure on MongoDB and allowed normal operations to resume.\n\n# **Root Cause**\n\nA higher-than-usual rate of concurrent updates against a frequently modified document used during pipeline execution produced a high rate of write conflicts in MongoDB. Each conflict triggered an automatic retry, which compounded under load and drove CPU utilization on the database to a sustained high level.\n\nAs CPU saturated:\n\n* MongoDB started queuing incoming commands instead of executing them immediately.\n* Operations dependent on those writes \u2014 pipeline step transitions and execution graph generation \u2014 slowed down or stalled while waiting for the queue to drain.\n\n# **Customer Impact**\n\nCustomers on Prod3 were impacted during two windows:\n\n* May 20, 2026: ~3:15 AM to ~4:45 AM PST \\(~1h 30m\\).\n* May 21, 2026: ~6:42 AM to ~7:47 AM PST \\(~1h 5m\\).\n\nIn both windows:\n\n* Pipeline executions ran significantly slower, with delays in step initialization and execution graph generation.\n* Pipelines with tight configured timeouts may have timed out as a result of the slowness.\n\nNo execution failures were observed beyond timeout-related effects, and no full outage occurred. All Harness modules that rely on Pipelines on Prod3 were affected for the duration of each incident.\n\n# **Resolution**\n\nAfter identifying that database contention was the bottleneck, we adjusted the pipeline service's database write settings to reduce overhead on each operation and added additional database capacity to absorb the load. Together these reduced write conflicts and allowed the queued commands to drain, after which pipeline operations returned to normal.\n\n# **Prevention and Improvements**\n\n1. Strengthen concurrency control around the high-contention update path so that concurrent updates to the same document no longer race in a way that produces write conflicts at scale.\n2. Strengthen load and concurrency testing on hot execution paths so similar contention patterns are caught before they affect production traffic.\n3. Improve proactive alerting on database write conflicts, queued commands, and sustained CPU utilization so the team can intervene before customer-facing latency is affected.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-05-21T09:34:39.165-07:00",
"resolved_inferred": false,
"started_at": "2026-05-21T06:55:48.842-07:00",
"state": "postmortem",
"title": "Prod3 pipeline execution UI slowness",
"updated_at": "2026-06-01T01:38:19.642-07:00",
"url": "https://stspg.io/fydt8rqlv14q"
},
{
"body": "# **Summary**\n\nOn May 20 and May 21, 2026, customers on the Prod3 cluster experienced degraded pipeline performance on two separate occasions. In both cases, pipeline executions ran slower than usual and execution graph rendering was delayed because database operations were being queued on the backing MongoDB instance. Both incidents were mitigated by adjusting database write settings and adding capacity headroom, which relieved pressure on MongoDB and allowed normal operations to resume.\n\n# **Root Cause**\n\nA higher-than-usual rate of concurrent updates against a frequently modified document used during pipeline execution produced a high rate of write conflicts in MongoDB. Each conflict triggered an automatic retry, which compounded under load and drove CPU utilization on the database to a sustained high level.\n\nAs CPU saturated:\n\n* MongoDB started queuing incoming commands instead of executing them immediately.\n* Operations dependent on those writes \u2014 pipeline step transitions and execution graph generation \u2014 slowed down or stalled while waiting for the queue to drain.\n\n# **Customer Impact**\n\nCustomers on Prod3 were impacted during two windows:\n\n* May 20, 2026: ~3:15 AM to ~4:45 AM PST \\(~1h 30m\\).\n* May 21, 2026: ~6:42 AM to ~7:47 AM PST \\(~1h 5m\\).\n\nIn both windows:\n\n* Pipeline executions ran significantly slower, with delays in step initialization and execution graph generation.\n* Pipelines with tight configured timeouts may have timed out as a result of the slowness.\n\nNo execution failures were observed beyond timeout-related effects, and no full outage occurred. All Harness modules that rely on Pipelines on Prod3 were affected for the duration of each incident.\n\n# **Resolution**\n\nAfter identifying that database contention was the bottleneck, we adjusted the pipeline service's database write settings to reduce overhead on each operation and added additional database capacity to absorb the load. Together these reduced write conflicts and allowed the queued commands to drain, after which pipeline operations returned to normal.\n\n# **Prevention and Improvements**\n\n1. Strengthen concurrency control around the high-contention update path so that concurrent updates to the same document no longer race in a way that produces write conflicts at scale.\n2. Strengthen load and concurrency testing on hot execution paths so similar contention patterns are caught before they affect production traffic.\n3. Improve proactive alerting on database write conflicts, queued commands, and sustained CPU utilization so the team can intervene before customer-facing latency is affected.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-05-20T05:38:18.000-07:00",
"resolved_inferred": false,
"started_at": "2026-05-20T03:48:29.082-07:00",
"state": "postmortem",
"title": "Prod3 Pipeline UI Updates on execution graph slowness",
"updated_at": "2026-06-01T01:36:57.420-07:00",
"url": "https://stspg.io/ygc9xh0hphnk"
},
{
"body": "# Summary\n\n Dashboards service in Prod3 experienced intermittent failures, causing dashboards to return errors and become unavailable to some customers. \n\n\u200c\n\n## Root Cause Analysis\n\nThe incident was caused by slow Looker `search_dashboards` API calls degrading from 7 seconds to over 30 seconds, which saturated the worker thread pool and prevented the health endpoint from responding to Kubernetes liveness and readiness probes.\n\n\u200c\n\n## Mitigation Steps Taken\n\n### Immediate Mitigations \n\n1. **Increased liveness probe timeout** from 15s to 30s \\(failure threshold kept at 3, allowing up to 90s tolerance\\)\n2. **Doubled  thread count** , increasing total concurrency\n3. **Increased pod replica count** to maintain higher minimum availability and reduce risk of all pods becoming simultaneously unavailable\n\n\u200c\n\n## Actions\n\nHarness will work on the following Action items to prevent recurrence.\n\n1. **Optimize Looker SDK timeout:** Tune timeout to Looker SDK `search_dashboards` calls to prevent indefinite thread occupation\n2. **Optimize Looker queries :** Investigate Looker-side query performance degradation to understand why `search_dashboards` latency increased",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-05-16T22:37:09.097-07:00",
"resolved_inferred": false,
"started_at": "2026-05-15T13:35:15.107-07:00",
"state": "postmortem",
"title": "Custom Dashboards are failing intermittently in prod3",
"updated_at": "2026-05-27T17:41:26.268-07:00",
"url": "https://stspg.io/t1l4225psm75"
},
{
"body": "same as [https://status.harness.io/incidents/pxl6nhjdnlwj](https://status.harness.io/incidents/pxl6nhjdnlwj)",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-05-21T09:34:51.147-07:00",
"resolved_inferred": false,
"started_at": "2026-05-14T21:54:59.591-07:00",
"state": "postmortem",
"title": "IACM Pipeline Struck In Prod4",
"updated_at": "2026-05-27T16:55:45.449-07:00",
"url": "https://stspg.io/lp4mp5076byv"
},
{
"body": "On May 14, 2026 , some customers running pipelines in the Prod4 production environment observed pipeline create and update requests that were slow or failed, and pipeline stages that did not start, produced no logs, and were eventually auto-aborted as \u201cstuck\u201d. \n\nThe issue was caused by an underlying compute node in our Prod4 cluster being recycled abruptly.\n\n## **Impact**\n\nDuring the incident window \\(approximately 5:38 PM PDT on May 14 to 9:47 PM PDT on May 14, 2026\\):\n\n* Some pipeline create and update requests on Prod4 were slow or failed.\n* Some Prod4 pipeline executions hung at the stage-start step, producing no logs, and were eventually auto-aborted as \u201cstuck\u201d after a timeout.\n* Behavior was intermittent \u2014 only pipelines whose requests were routed to an affected service pod were impacted; other pipelines continued to execute normally.\n\nThere was **no data loss**. The majority of pipelines on Prod4 continued to execute successfully throughout the incident \u2014 the primary impact was that affected create/update requests slowed down or failed, and a subset of pipelines could not progress and had to be aborted and re-run after mitigation. Overall service availability was degraded during this window.\n\n## **Root Cause**\n\nDuring the incident, an underlying compute node in our Prod4 cluster was recycled by the cloud provider without completing its normal graceful-drain process, so the supporting-service pods running on that node were terminated abruptly. As a result, in-flight requests from the backend service to those pods were left without a response. \n\n## **Mitigation**\n\nHarness completed the following immediate mitigation steps:\n\n* Restarted the affected supporting-service pods to restore healthy targets.\n* Restarted the pipeline service in Prod4 to clear the blocked worker threads. This is what fully restored normal pipeline create/update behavior and stage-start behavior; restarting only the supporting service was not enough on its own.\n* Confirmed pipeline executions returned to normal and updated the status page to mitigated.\n\nThese actions restored pipeline execution behavior and resolved the customer-facing impact.\n\n## **Action Items**\n\nTo reduce the risk of recurrence and improve detection, the following actions are in various stages of being implemented:\n\n* Optimize timeouts to the pipeline service\u2019s plan-creation requests so that when a supporting service goes away unexpectedly, the worker threads recover automatically instead of remaining blocked.\n* Investigate the abrupt node-recycle behavior in Prod4 with our cloud provider to ensure pods running on a recycled node receive a graceful shutdown signal in the future.\n* Add proactive paging alerts on service worker-thread saturation, so this failure mode is detected before it becomes a impacting issue",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-05-14T07:10:52.072-07:00",
"resolved_inferred": false,
"started_at": "2026-05-14T06:36:21.564-07:00",
"state": "postmortem",
"title": "Pipeline Updates in Prod4 is taking time.",
"updated_at": "2026-05-27T16:50:51.474-07:00",
"url": "https://stspg.io/67q50lr5r5qh"
},
{
"body": "## **Summary**\n\nOn May 14, 2026 , some customers running pipelines in the Prod4 production environment observed pipeline create and update requests that were slow or failed, and pipeline stages that did not start, produced no logs, and were eventually auto-aborted as \u201cstuck\u201d. \n\nThe issue was caused by an underlying compute node in our Prod4 cluster being recycled abruptly. As a downstream side effect, internal worker threads in the pipeline service became blocked waiting on responses that would never arrive, which caused some pipeline create/update requests and stage starts to slow down or fail.\n\n## **Impact**\n\nDuring the incident window \\(approximately 5:38 PM PDT on May 14 to 9:47 PM PDT\\):\n\n* Some pipeline create and update requests on Prod4 were slow or failed.\n* Behavior was intermittent \u2014 only pipelines whose requests were routed to an affected service pod were impacted; other pipelines continued to execute normally.\n\nThere was **no data loss**. The majority of pipelines on Prod4 continued to execute successfully throughout the incident \u2014 the primary impact was that affected create/update requests slowed down or failed, and a subset of pipelines could not progress and had to be aborted and re-run after mitigation. Overall service availability was degraded during this window.\n\n## **Root Cause**\n\nThe pipeline service coordinates pipeline plan creation by sending requests to several internal supporting services. \n\nDuring the incident, an underlying compute node in our Prod4 cluster was recycled by the cloud provider without completing its normal graceful-drain process, so the supporting-service pods running on that node were terminated abruptly. As a result, in-flight requests from the pipeline service to those pods were left without a response. \n\n## **Mitigation**\n\nHarness completed the following immediate mitigation steps:\n\n* Restarted the affected supporting-service pods to restore healthy targets.\n* Restarted the pipeline service in Prod4 to clear the blocked worker threads. \n* Confirmed pipeline executions returned to normal and updated the status page to mitigated.\n\nThese actions restored pipeline execution behavior and resolved the customer-facing impact.\n\n## **Action Items**\n\nTo reduce the risk of recurrence and improve detection, the following actions are in various stages of being implemented:\n\n* Enhance timeout configuration to the pipeline service\u2019s plan-creation requests so that when a supporting service goes away unexpectedly, the worker threads recover automatically instead of remaining blocked.\n* Add per-target instrumentation on the pipeline service\u2019s plan-creation request fanout \\(request count, latency, in-flight requests\\) so the affected supporting service can be identified as soon as possible during an incident.\n* Investigate the abrupt node-recycle behavior in Prod4 with our cloud provider to ensure pods running on a recycled node receive a graceful shutdown signal in the future.\n* Add proactive paging alerts on pipeline-service worker-thread saturation, so this failure mode is detected before it becomes customer-visible.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-05-14T08:02:36.406-07:00",
"resolved_inferred": false,
"started_at": "2026-05-14T06:00:24.000-07:00",
"state": "postmortem",
"title": "IACM Pipeline Struck In Prod4",
"updated_at": "2026-05-22T09:35:25.936-07:00",
"url": "https://stspg.io/w1sx5thrncf6"
},
{
"body": "## Summary\n\nBetween May 17 and May 19, 2026, customers in the APAC region experienced intermittent unresponsiveness with the Harness platform. The incident occurred twice within a two-day window. In both cases, the GraphQL service became degraded, causing requests to fail or time out for affected users. The service was restored each time by restarting the impacted pods. A permanent alerting fix is in progress and expected to reach production shortly.\n\n\u200c\n\n## Impact\n\n* APAC-region customers experienced intermittent inability to access the Harness platform UI and API\n* No data loss was reported; the impact was limited to service availability\n\n\u200c\n\n## Root Cause\n\nOne of the two service pods in the APAC cluster entered a degraded state where downstream calls were timing out. While the downstream services showed no visible increase in response time on their end, the pod was unable to receive responses from them in time likely due to a network-level communication issue between the pod and its downstream dependencies.\n\n\u200c\n\n## Mitigation\n\nIdentified and restarted problematic nodes. Recovery was confirmed by the engineering team before the incident was marked resolved.\n\n## Next Steps\n\n**Improved Alerting and Proactive Detection** \u2014 Implemented new alerts for elevated response times to ensure any recurrence is detected and mitigated before impacting customers.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-05-14T04:13:10.796-07:00",
"resolved_inferred": false,
"started_at": "2026-05-14T03:57:39.814-07:00",
"state": "postmortem",
"title": "Traceable APAC Cluster Slowness",
"updated_at": "2026-05-22T13:37:21.158-07:00",
"url": "https://stspg.io/7qypzpssnnq9"
},
{
"body": "## Summary\n\nOn May 12, 2026, Harness's Secure Connect service experienced a significant disruption that affected customers using Secure Connect for git connector checks and CI pipeline execution. Affected customers experienced connection timeouts and failures when attempting to use Secure Connect-enabled connectors, while connectors operating without Secure Connect continued to function normally.\n\nThe disruption was caused by an automated infrastructure maintenance event in our Google Kubernetes Engine \\(GKE\\) environment that left a critical load balancer in an unreachable state. The issue was resolved by recreating the affected load balancer and updating the corresponding DNS record to restore connectivity.\n\n**Impact:**\n\n* Secure Connect-enabled git connector tests failing with connection timeouts\n* CI pipelines hanging or failing when Secure Connect was enabled\n* Connectors operating without Secure Connect were unaffected\n\n## Root Cause\n\nThe disruption was triggered by an automated GKE \\(Google Kubernetes Engine\\) control plane upgrade on the cluster hosting our internal Secure Connect FRPS \\(Fast Reverse Proxy Server\\) infrastructure.\n\nThe upgrade executed in two sequential phases in rapid succession. Each phase restarts the GKE cloud controller manager, which is responsible for managing the lifecycle of Kubernetes LoadBalancer service IP addresses. As part of its normal reconciliation process, the controller performs a delete-then-re-insert cycle on IP address reservations when it restarts.\n\nThe second phase of the upgrade began before the first reconciliation cycle had fully completed. This created a race condition where the controller deleted the IP address reservation for the internal FRPS load balancer but was interrupted before it could re-acquire it.  No Harness application deployments caused or contributed to the issue.\n\n\u200c\n\n## Mitigation\n\nThe following steps were taken to restore service:\n\n1. **FRPS pod restart** \u2014 An initial rollout restart of the FRPS deployment was performed. This restored client connectivity to the FRPS server but did not resolve the underlying IP orphaning issue.\n2. **Load balancer recreation** \u2014 The internal Kubernetes service \\(`frps-internal`\\) was deleted and recreated, triggering GKE to provision a new load balancer with a new IP address \n3. **DNS record update** \u2014 The Cloud DNS A record for  was updated to point to the new load balancer IP. This restored end-to-end routing for all Secure Connect traffic.\n\nFull connectivity was confirmed via internal testing and subsequently verified by affected customers.\n\n## Preventative Actions\n\nHarness is committed to preventing this class of incident from recurring. The following actions are being implemented:",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-05-12T16:10:39.274-07:00",
"resolved_inferred": false,
"started_at": "2026-05-12T14:38:03.677-07:00",
"state": "postmortem",
"title": "Hosted CI customers using secure connect are facing connectivity issues",
"updated_at": "2026-05-19T19:04:00.533-07:00",
"url": "https://stspg.io/mw9ds4fty03n"
},
{
"body": "## **Summary**\n\nBetween 10:07 AM\u201310:39 AM PST on Tuesday, May 12, 2026, customers using the prod1, prod2, and prod3 Production clusters experienced elevated latency and intermittent service degradation. During this timeframe, customers observed delegate timeouts, login failures, and pipeline execution failures\n\n## **Root Cause**\n\nA recently introduced configuration change to a common infrastructure component caused unexpected resource pressure across nodes in the prod1, prod2, and prod3 production clusters. The peak traffic exacerbated the resource utilization and introduced elevated latency across several critical platform services.\n\n## **Impact**\n\n1. Customers in prod1, prod2, and prod3 experienced login and access failures for approximately 20 minutes.\n2. Delegate connectivity was intermittently impacted during the incident window.\n3. Pipeline executions and API requests experienced elevated failure rates and latency during the incident window.\n\n## **Remediation**\n\n* Immediately Rolled back to the previous stable release, restoring customer pipeline functionality and alleviating node pressure.\n\n## **Action Items**\n\n1. Enhance perf testing to include such workloads so that we can catch issues before we hit production.\n2. Increase capacity across clusters to make sure we have enough headroom to absorb the traffic surges.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-05-12T10:55:04.760-07:00",
"resolved_inferred": false,
"started_at": "2026-05-12T10:33:31.277-07:00",
"state": "postmortem",
"title": "Platform access issues in Prod1/Prod2/Prod3",
"updated_at": "2026-05-12T17:03:58.881-07:00",
"url": "https://stspg.io/z2z0n9078zlt"
},
{
"body": "### Summary\n\nOn May 12, 2026, Harness identified an issue affecting Kubernetes-based CI builds in the prod2 environment. The issue was caused by the unintended enablement of an internal feature flag associated with a CI performance optimization.\n\nEffect was seen for build steps using older `ci-1.16.81` CI addon version \\(published 1 year ago\\). Affected builds using older CI addon versions experienced the following behavior during the incident window:\n\n* Secrets referenced in Run step commands resolved as empty values, causing build failures\n\nNo customer action is required.\n\n### Impact\n\n* Affected Environment: prod2\n* Affected Builds\n\n    * Kubernetes CI builds executed during the incident window\n    * Builds using CI addon versions older than `ci-1.16.81`\n    \n* Customer Impact:\u00a0\n\n    * Secret expressions resolving as empty strings during execution\n    \n\n\u200c\n\n### Root Cause\n\nHarness introduced a CI optimization feature intended to improve Kubernetes pipeline execution performance. The feature depended on functionality available only in newer CI addon versions \\(`ci-1.16.81` and later\\).\n\nOn May 12, 2026, the feature flag controlling this optimization was unintentionally enabled for customer accounts in the prod2 environment because of misconfigured  checks. The intent of change was to verify in certain internal accounts as part of a phased roll out of this feature.\n\n\u200c\n\nFor environments running older addon versions:\n\n\u200c\n\n* Runtime secret resolution did not execute correctly, causing secrets referenced in Run step commands to resolve as empty values\n\n\u200c\n\nThe issue was mitigated immediately by disabling the feature flag globally.\n\n### Mitigation Actions\n\nHarness completed the following remediation steps:\n\n\u200c\n\n* Disabled the feature flag globally\n* Verified successful pipeline execution after rollback\n* Identified and deleted all affected downloadable log entries from GCP storage\n* Sent communications to affected customers\n\n### Preventive Actions\n\nHarness is implementing the following safeguards to prevent recurrence:\n\n* Enforcing addon version compatibility validation before enabling feature flags that depend on addon functionality\n* Documenting compatibility requirements between CI manager and addon versions",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-05-12T07:35:54.000-07:00",
"resolved_inferred": false,
"started_at": "2026-05-12T02:50:29.000-07:00",
"state": "postmortem",
"title": "Secret resolution failures in CI executions - PROD2",
"updated_at": "2026-05-19T18:42:01.395-07:00",
"url": "https://stspg.io/dcyn647w5x8b"
},
{
"body": "### **Summary**\n\nOn May 7, 2026, customers using the Vercel integration experienced delays in synchronizing Feature Management & Experimentation \\(FME\\) changes to Vercel. The issue was caused by a dependency update in the integration service that prevented outbound synchronization requests from being authenticated correctly. The issue was resolved on May 11, 2026 through a service fix and validation of synchronization workflows. \u00a0\n\n### **Impact**\n\n* Customers using the Vercel integration experienced delays propagating FME changes to Vercel\n* Customers observed stale configuration data in Vercel after updating changes in FME\n* Other FME functionality remained operational\n* No data loss occurred\n\n### **Root Cause**\n\nA deployment to the `integrations-service-v2` service introduced dependency updates that removed a required AWS SDK component used during request signing and authentication for synchronization calls.\n\nAs a result, outbound requests from the FME integration service to the synchronization endpoint could not be authenticated correctly, causing synchronization operations to fail.\n\nThe issue was not immediately detected because:\n\n* Existing monitoring metrics did not have alerting configured\n* The failure mode did not surface as a critical application startup error\n* Integration-level automated test coverage for this synchronization workflow was insufficient\n\n### **Resolution**\n\nThe missing dependency was restored and the synchronization workflow was validated across affected integrations. Monitoring and validation checks confirmed successful synchronization behavior after deployment.\n\n### **Preventive Actions**\n\nTo prevent such issues from happening again, We are implementing the following improvements:\n\n* Enhance end-to-end integration tests for the Vercel synchronization workflow\n* Enhance our proactive alerts for synchronization failures and elevated error rates",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-15T12:28:18Z",
"resolved_at": "2026-05-11T16:15:48.044-07:00",
"resolved_inferred": false,
"started_at": "2026-05-11T15:42:28.505-07:00",
"state": "postmortem",
"title": "Vercel Integration for FME is having delays to synchronize changes",
"updated_at": "2026-05-12T18:55:44.394-07:00",
"url": "https://stspg.io/d8gxhv2m69sz"
},
{
"body": ".",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-15T12:28:18Z",
"resolved_at": "2026-05-11T00:16:45.588-07:00",
"resolved_inferred": false,
"started_at": "2026-05-10T23:48:03.469-07:00",
"state": "postmortem",
"title": "Delegate Upgrades are failing for some customers",
"updated_at": "2026-06-23T16:51:05.375-07:00",
"url": "https://stspg.io/lhxhkmn3tjvz"
},
{
"body": "On May 8, 2026, a CI performance optimization feature flag was inadvertently impacted customer with older delegate version. The feature was designed to improve CI pipeline initialization performance by deferring secret resolution to runtime on supported delegate versions.\n\nCustomers running delegate versions that did not support the new runtime secret resolution capability experienced pipeline execution failures where secrets referenced in Run step commands resolved as empty strings.\n\nThe issue primarily impacted customers running older delegate versions outside the currently supported delegate release window. Customers on supported delegate versions were not impacted.\n\nThe issue was mitigated by disabling the feature flag, which immediately restored the prior secret resolution behavior. No customer configuration changes were required.\n\n**Impact**\n\n* Pipeline execution failures for affected Kubernetes CI builds\n* Secrets referenced in Run step commands resolved as empty strings\n* Customers running supported delegate versions were not affected\n\n**Root Cause**\n\nA feature flag intended for a limited internal rollout was unintentionally enabled for a broader set of accounts during a routine rollout process. The feature relied on delegate-side functionality available only in newer delegate versions. Older delegate versions did not support the required runtime secret resolution behavior, resulting in pipeline failures.\n\n\u200c\n\n**Resolution**\n\nThe feature flag was disabled globally, restoring the previous behavior where secrets are resolved during pipeline initialization. Impacted pipelines recovered immediately after rollback.\n\n**Preventive Actions**\n\nHarness is implementing the following improvements:\n\n* Enhancements to the feature flag approval workflow to more clearly display rollout scope.\n* Improved process and documentation for feature flags that depend on minimum delegate versions\n* Expanded monitoring and alerting for secret resolution failures in CI pipelines",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-14T12:29:13Z",
"resolved_at": "2026-05-08T11:10:00.000-07:00",
"resolved_inferred": false,
"started_at": "2026-05-08T11:10:00.000-07:00",
"state": "postmortem",
"title": "Pipeline and secret resolution failures",
"updated_at": "2026-05-18T17:35:16.119-07:00",
"url": "https://stspg.io/0lv9hd2cb6rg"
},
{
"body": "### Summary\n\nOn May 8, 2026, between approximately 17:00 and 18:37 UTC, a small fraction of SDK requests to [sdk.split.io](http://sdk.split.io/) returned errors or experienced elevated latency. The cache that fronts [sdk.split.io](http://sdk.split.io/) continued to serve the vast majority of requests normally throughout the event. The issue was mitigated by increasing backend capacity, and all systems returned to baseline by 18:37 UTC.\n\n### Root Cause\n\n[sdk.split.io](http://sdk.split.io/) is served by a CDN that caches flag data and forwards a small percentage of requests to a backend service when cached data is not available. During this event, an elevated rate of cache updates caused the CDN to forward roughly 40x its normal volume of requests to the backend. Sustained over several hours, this exceeded the backend's available capacity, leading to elevated latency, increased error rates on requests forwarded to the backend, and intermittent service restarts during recovery.\n\n### Impact\n\n* The CDN continued to serve 97.5% of requests directly from cache as normal across the affected window.\n* The remaining requests forwarded to the backend saw elevated errors \\(approximately 0.35% of total SDK traffic during the window\\) or latency until backend capacity was restored.\n* SDKs continued to evaluate flags normally using their locally cached flag data; no flag evaluation correctness issues occurred. New SDK instances initializing and getting an error retried and/or signaled the timeout depending on configurations.\n* No data loss occurred.\n\n### Remediation\n\n* Increased backend capacity across multiple dimensions, including service replicas, compute and memory per replica, and database connection capacity.\n* Service stabilized at the increased capacity profile after the changes were applied.\n\n### Action Items\n\n* Allow additional backend capacity parameters to be tuned at runtime for quicker scaling response.\n* Add monitoring and alerting on identified leading indicators for earlier detection of saturation trends.\n* Adopt more responsive autoscaling for the backend so it can react to traffic surges more aggressively than the current autoscaler.\n* Improve load-shedding so the backend drops excess requests cleanly under saturation rather than queueing to failure.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-14T12:29:13Z",
"resolved_at": "2026-05-08T11:53:07.314-07:00",
"resolved_inferred": false,
"started_at": "2026-05-08T10:05:39.129-07:00",
"state": "postmortem",
"title": "FME SDK are experiencing elevated error rates and delays in response.",
"updated_at": "2026-05-18T10:27:13.035-07:00",
"url": "https://stspg.io/3k9hf2mljyds"
},
{
"body": "Dupe of [Deployment Degradation \u2013 Failures in Prod1,2,3](https://manage.statuspage.io/pages/zmnt6tkys0q0/incidents/v4nxhkv1mvdg)",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "critical",
"last_seen": "2026-09-08T12:29:58Z",
"resolved_at": "2026-05-07T05:54:51.838-07:00",
"resolved_inferred": false,
"started_at": "2026-05-07T01:00:28.000-07:00",
"state": "postmortem",
"title": "Deployment Degradation - Slowness in Prod3",
"updated_at": "2026-05-19T17:47:42.798-07:00",
"url": "https://stspg.io/0wdy1s4bqpnv"
},
{
"body": "### Incident Summary\n\n  \nOn May 6 at 11:50 PM PST, we deployed a configuration change to one of our core pipeline  \nservices. This change introduced an unintended interaction with our database layer, causing a  \nsignificant increase in write load. The resulting pressure degraded query and command  \nthroughput across the platform\n\n### Root Cause\n\n  \nThe configuration change introduced a blocking condition on expression evaluation in the  \npipeline service. When executions encountered blocked expressions, they failed and retried  \nrepeatedly, generating a write storm against the database and the  throughput went 4x\n\n\u200c\n\n### Remediation \n\n  \n\u25cf Rolled back the configuration change.  \n\u25cf Applied database-level tuning to reduce write pressure and accelerate backlog drainage  \n\u25cf Performed a controlled failover to a healthy database node to restore throughput  \n\u25cf Scaled up database nodes to provide sufficient capacity for full recovery\n\n\u200c\n\n### Preventive Actions\n\nTo prevent from such Issues happening again, we are focussing on:\n\n  \n1\\. Increased database resilience: We are implementing automated load-shedding thresholds  \nthat trigger on leading indicators \\(replication lag, session depth, op latency\\) before the database  \nreaches saturation, preventing retry storms from compounding into full degradation events.  \n\n2\\. We are optimizing databases so that we can increase write throughput by an order of magnitude and enable independent scaling of customer data workloads. This would have allowed us to drain the message backlog nearly instantaneously during this incident.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-05T08:54:55Z",
"resolved_at": "2026-05-07T04:23:53.420-07:00",
"resolved_inferred": false,
"started_at": "2026-05-07T00:54:42.000-07:00",
"state": "postmortem",
"title": "Deployment Degradation \u2013 Failures in Prod1,2,3",
"updated_at": "2026-05-19T17:46:55.987-07:00",
"url": "https://stspg.io/d9dlxwt58060"
}
]
}