{
"vendor": "Flexera",
"slug": "flexera",
"platform": "statuspage",
"status_url": "https://status.flexera.com",
"last_checked": "2026-09-16T12:28:20Z",
"last_state": "degraded",
"history_backfilled": true,
"first_watched": "2026-09-04T07:06:16Z",
"incidents": [
{
"body": "Incident Description: We are currently experiencing an issue affecting the Flexera Community site. Due to an ongoing outage impacting a third-party service provider, customers may be unable to access the Flexera Community portal or may experience intermittent and unreliable behaviour. \nAdditionally, customers may be unable to create, update, or interact with support cases through the Community site. Emails sent to support@flexera.com during this period may not be received or processed until the underlying outage has been resolved.\nThis third party service provider outage also impacts the proactive monitoring ability for Spot support teams to detect any customer environment issues.\n\nPriority: P3\n\nRestoration Activity: Our technical teams are actively monitoring the situation and working with the third-party provider to track restoration efforts. We will continue to assess customer impact and provide further updates as additional information becomes available.",
"first_seen": "2026-09-16T12:28:20Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"started_at": "2026-09-16T02:04:03.771-07:00",
"state": "identified",
"title": "Flexera Community- Service unavailable",
"updated_at": "2026-09-16T02:38:58.553-07:00",
"url": "https://stspg.io/2j5w0n3hlhwj"
},
{
"body": "Our teams have continued to monitor the service closely, and no new issues have been observed. The backlog of queued work that accumulated during the disruption is processing successfully and continuing to decrease as expected.\nBased on the stability observed since recovery, this incident is now resolved. We will continue to monitor processing to ensure the remaining backlog clears completely and the service continues to operate as expected.",
"first_seen": "2026-09-16T12:28:20Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-09-16T01:24:43.429-07:00",
"resolved_inferred": false,
"started_at": "2026-09-15T16:05:35.423-07:00",
"state": "resolved",
"title": "Flexera -  EU - Policies/Automation Unavailable",
"updated_at": "2026-09-16T01:24:43.444-07:00",
"url": "https://stspg.io/9cfznx9gz20j"
},
{
"body": "Our investigation identified a connectivity issue within a service dependency path that affected customer access in the North America environment. Our technical teams updated the routing configuration, successfully restoring connectivity between the affected services. Service has been fully restored and is operating normally. The issue is resolved.",
"first_seen": "2026-09-15T12:28:18Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-09-14T16:00:09.531-07:00",
"resolved_inferred": false,
"started_at": "2026-09-14T15:23:01.874-07:00",
"state": "resolved",
"title": "Flexera One - NAM - Access Disruption",
"updated_at": "2026-09-14T16:00:09.548-07:00",
"url": "https://stspg.io/0kdlqjm7rw6y"
},
{
"body": "Incident Description: Our teams earlier detected an intermittent issue that affected a subset of customers using Flexera One IT Asset Management in the EU environment. Affected customers may have intermittently experienced missing ITAM menu items, blank or pink error screens, or difficulty loading ITAM pages, while other Flexera One modules remained accessible.\n\nPriority: P2\n\nImpact Start: September 10, 2026 , 2:00 AM PDT\nImpact End: September 10, 2026 , 6:00 AM PDT\n\nResolution: The issue self-resolved, and affected customers regained access to ITAM functionality. Our technical teams monitored the environment and confirmed that error levels decreased and the service returned to a stable state.\n\nNext Actions: Our teams will continue to investigate the underlying root cause for the issue and further details will be shared via a post mortem report.",
"first_seen": "2026-09-12T12:30:15Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-09-12T04:40:00.204-07:00",
"resolved_inferred": false,
"started_at": "2026-09-12T04:40:00.145-07:00",
"state": "resolved",
"title": "Flexera One - IT asset management- EU - Intermittent Black/Pink screen error",
"updated_at": "2026-09-12T04:40:09.007-07:00",
"url": "https://stspg.io/605jl64jldbp"
},
{
"body": "Our investigation remains ongoing, and recent testing has shown that the previously deployed hotfix has not fully resolved the issue.\n\nTo address this, our teams have developed and deployed an updated hotfix to a validation environment, where additional testing and verification activities are currently underway. The investigation has narrowed the scope of the issue, and teams are actively working to identify the remaining factors contributing to the behaviour.\n\nWe remain focused on implementing a permanent resolution and will continue to provide updates as progress is made.",
"first_seen": "2026-09-11T12:29:42Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"started_at": "2026-09-11T00:47:22.748-07:00",
"state": "identified",
"title": "Flexera One - IT Asset Management - APAC & EU - Inventory .ZIP File Processing Delays",
"updated_at": "2026-09-16T01:15:46.766-07:00",
"url": "https://stspg.io/zmmg3mrj8myh"
},
{
"body": "The hotfix was successfully deployed, and our teams closely monitored the services overnight to validate the effectiveness of the remediation and ensure continued stability.\n\nNo new errors were observed following the deployment, and the affected services remained stable throughout the monitoring period. Based on these results, this incident is now considered resolved.\n\nOur teams will continue to investigate the underlying cause and identify any additional improvement opportunities. Further details, including the root cause analysis and corrective actions, will be shared in a post-mortem report.",
"first_seen": "2026-09-09T12:30:34Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-09-15T22:28:52.223-07:00",
"resolved_inferred": false,
"started_at": "2026-09-08T20:40:42.384-07:00",
"state": "resolved",
"title": "Flexera One \u2013 IT Asset Management \u2013 EU \u2013 Reconciliation Failures",
"updated_at": "2026-09-15T22:28:52.239-07:00",
"url": "https://stspg.io/3q38jfcj7h35"
},
{
"body": "Post-deployment validation activities have concluded, and service availability and platform stability have been confirmed. Access to Flexera One IT Asset Management services in the APAC region has been fully restored.",
"first_seen": "2026-09-09T12:30:34Z",
"impact": "critical",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-09-08T13:57:58.739-07:00",
"resolved_inferred": false,
"started_at": "2026-09-08T11:46:47.668-07:00",
"state": "resolved",
"title": "Flexera One - IT Asset Management - APAC - Service Unavailable",
"updated_at": "2026-09-08T13:57:58.754-07:00",
"url": "https://stspg.io/zj8b8chwd3h6"
},
{
"body": "Scheduled maintenance is currently in progress. We will provide updates as necessary.",
"first_seen": "2026-09-06T12:28:17Z",
"impact": "maintenance",
"last_seen": "2026-09-06T12:28:17Z",
"resolved_at": "2026-09-06T15:32:04Z",
"resolved_inferred": true,
"started_at": "2026-09-05T07:00:00.000-07:00",
"state": "maintenance",
"title": "Scheduled Maintenance: Flexera One - IT Visibility - NAM",
"updated_at": "2026-09-05T07:00:11.999-07:00",
"url": "https://stspg.io/tq2xnn59r0qr"
},
{
"body": "Scheduled maintenance is currently in progress. We will provide updates as necessary.",
"first_seen": "2026-09-05T05:19:27Z",
"impact": "maintenance",
"last_seen": "2026-09-05T12:15:31Z",
"resolved_at": "2026-09-06T12:28:17Z",
"resolved_inferred": true,
"started_at": "2026-09-04T21:00:00.000-07:00",
"state": "maintenance",
"title": "Scheduled Maintenance: Flexera One - IT Visibility - EU",
"updated_at": "2026-09-04T21:00:07.309-07:00",
"url": "https://stspg.io/kgdbvgy7wfxk"
},
{
"body": "Scheduled maintenance is currently in progress. We will provide updates as necessary.",
"first_seen": "2026-09-05T05:19:27Z",
"impact": "maintenance",
"last_seen": "2026-09-05T12:15:31Z",
"resolved_at": "2026-09-06T12:28:17Z",
"resolved_inferred": true,
"started_at": "2026-09-04T19:00:00.000-07:00",
"state": "maintenance",
"title": "Scheduled Maintenance: Flexera One - IT Visibility - APAC",
"updated_at": "2026-09-04T19:00:07.598-07:00",
"url": "https://stspg.io/kwnv96fjhf56"
},
{
"body": "**Description:** Spot Ocean - Resource State Discrepancies and Under-Capacity Alerts\n\n**Timeframe:** August 25, 2026, 6:15 PM PDT \u2013 August 25, 2026, 10:22 PM PDT  \n\n**Incident Summary**\n\nOn Tuesday, August 25, 2026, at 6:15 PM PDT, Ocean customers began experiencing service degradation affecting resource state visibility and management operations.  \n  \nDuring the incident, customers may have observed instances remaining in a Resuming state longer than expected, under-capacity alerts, and discrepancies between resource status information displayed in Spot and AWS. In some cases, resource state information displayed within Spot did not accurately reflect the actual state of resources in AWS.  \n  \nTechnical teams investigated the issue and identified degradation within a central Ocean service responsible for processing resource monitoring and management requests. As a result, monitoring and management operations slowed or stopped, resulting in broader service degradation across Ocean.  \n  \nThe affected service was restarted, restoring normal request processing and service functionality. Following validation and monitoring activities, service stability was confirmed and the incident was resolved.  \n\n**Root Cause**\n\nInvestigation determined that a service responsible for processing resource monitoring and management requests experienced prolonged delays while communicating with a dependent service. As those delays accumulated, the service became unable to effectively process new requests.  \n  \nAlthough the service continued appearing operational, it was no longer able to process requests normally. This resulted in delayed resource state updates and contributed to the customer-facing symptoms observed during the incident, including resource state discrepancies, under-capacity alerts, and instances remaining in a Resuming state longer than expected.\n\n  \n**Remediation Actions**\n\nThe following actions were taken during the incident response:\n\n1. **Incident Investigation Initiated:** Technical teams responded to production alerts and customer reports to assess the scope and impact of the issue. \n2. **Service Degradation Identified: I**nvestigation determined that the affected service was no longer able to effectively process new monitoring and resource management requests. \n3. **Service Recovery Performed:** The affected service components were restarted, restoring normal request processing and service functionality. \n4. **Post-Recovery Validation Completed:** Technical teams monitored service behavior and validated stability before declaring the incident resolved.\n\n  \n**Future Preventative Measures**\n\nTechnical teams have initiated follow-up work to reduce the likelihood of similar issues occurring in the future.\n\n1. **Monitor Service Reliability Improvements:** Technical teams have identified an existing service behavior that contributed to this incident and are implementing improvements to enhance overall service reliability.\n2. **Service Design and Configuration Review:** Technical teams are reviewing the service design, configuration, and request processing behavior involved in this incident to identify additional opportunities to improve resiliency.\n3. **Additional Investigation and Corrective Actions:** Technical teams are continuing to assess the findings identified during the investigation and will implement any additional corrective actions determined to be relevant to this incident.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-08-25T22:57:56.232-07:00",
"resolved_inferred": false,
"started_at": "2026-08-25T20:44:32.018-07:00",
"state": "postmortem",
"title": "Spot - Service Degradation",
"updated_at": "2026-09-02T13:42:45.938-07:00",
"url": "https://stspg.io/yjfyh10vnd8v"
},
{
"body": "**Description:** Flexera One - IT Asset Management - North America - Inventory Processing Delays\n\n**Timeframe:** August 24, 2026, at 5:33 AM PDT to August 27, 2026, at 9:30 AM PDT\n\n**Incident Summary**\n\nOn Monday, August 24, 2026, at 5:33 AM PDT, reports surfaced of delayed inventory processing affecting some Flexera One IT Asset Management customers in the North America region. Customer inventory files continued to be uploaded, but some files were not progressing through downstream processing as expected. This resulted in growing processing backlogs and delayed availability of updated inventory data for affected customers. The issue affected inventory processing rather than access to the IT Asset Management platform\n\nInvestigation identified that the issue followed a platform upgrade in the North America environment. During the upgrade activity, inventory upload requests experienced an elevated number of service unavailable responses, connection resets, and other connection-related failures. These conditions caused a cascade of failures within an inventory processing component. Although most inventory processing streams recovered automatically after the affected service components restarted, some streams remained stalled with inventory files queued for processing.\n\nTechnical teams restarted the affected service components, which restored processing for the impacted workflows, although the accumulated inventory backlog initially cleared more slowly than expected. Processing capacity was increased, and workload, concurrency, and throughput settings were adjusted to improve processing performance. Service health checks were also updated as part of the stabilization activity. These actions increased processing throughput and enabled the accumulated backlogs to be worked through.\n\nOn Thursday, August 27, 2026, at 9:30 AM PDT, the incident was considered resolved after the implemented measures restored stable inventory processing and the remaining production backlogs continued to decrease. Technical teams continued monitoring processing performance and backlog recovery following resolution. Subsequent validation confirmed that production inventory backlogs had cleared, with only isolated non-production testing backlogs remaining.\n\n**Root Cause**\n\n**Primary Root Cause:**\n\nFollowing a platform upgrade in the North America environment, inventory processing workflows experienced elevated service unavailable responses, connection interruptions, and related processing failures. While most processing activity recovered automatically, a subset of inventory processing workflows did not fully recover, resulting in inventory files remaining queued and unprocessed. This led to inventory processing backlogs and delayed inventory updates for affected customers.\n\n**Contributing Factors:**    \n\u2022\tSome inventory processing workflows remained stalled following the initial service disruption, preventing queued inventory data from being processed and contributing to backlog growth.    \n\u2022\tExisting monitoring capabilities did not provide visibility into this specific processing condition, resulting in abnormal backlog growth being identified through investigation rather than automated alerting.\n\n**Remediation Actions**\n\nThe following actions were taken during the incident response:\n\n1. **Incident Investigation Initiated:** Technical teams investigated reports of delayed inventory processing, reviewed processing behavior, and validated the scope of impact affecting inventory processing workflows in the North America environment.\n2. **Service Recovery Actions Performed:** Affected service components were restarted to restore inventory processing activity and recover impacted processing workflows.\n3. **Processing Capacity Increased:** Additional processing capacity was introduced to improve throughput, reduce processing backlogs, and support backlog recovery across affected inventory processing workflows.\n4. **Processing Optimizations Applied:** Technical teams adjusted throughput, workload, and concurrency settings to improve inventory processing performance and accelerate backlog reduction.\n5. **Stability and Recovery Validation Performed:** Processing throughput, backlog reduction, service health, and application stability were continuously monitored to verify successful recovery and confirm that inventory processing had returned to normal operating levels.\n\n**Future Preventative Measures**\n\nFollowing this incident, technical teams continue to assess the conditions that contributed to the inventory processing disruption and are implementing improvements to strengthen inventory processing reliability, detection, and recovery capabilities.\n\n1. **Validation and Monitoring Following Platform Changes:** Technical teams are reviewing validation and monitoring activities associated with platform changes to improve the early identification of unexpected inventory processing behavior following maintenance and upgrade activities.\n2. **Enhanced Inventory Processing Monitoring and Alerting:** Additional monitoring and alerting will be implemented to provide earlier visibility into abnormal inventory backlog growth and inventory processing disruptions.\n3. **Inventory Processing Recovery Improvements:** Improvements will be implemented to strengthen inventory processing recovery behavior and reduce the likelihood of processing workflows remaining stalled following unexpected service interruptions.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-08-27T09:38:30.951-07:00",
"resolved_inferred": false,
"started_at": "2026-08-24T06:04:05.058-07:00",
"state": "postmortem",
"title": "Flexera One - IT Asset Management - North America - Inventory Processing Delays",
"updated_at": "2026-09-04T11:39:56.626-07:00",
"url": "https://stspg.io/y5c0j39bh3xm"
},
{
"body": "**Description:** Flexera One - IT Asset management - APAC - Access disruption\n\n**Timeframe:**  August 18, 2026, 2:10 AM PDT to August 18, 2026, 4:26:03 AM PDT\n\n\u200c\n\n**Incident Summary**\n\n\u200c\n\nOn Tuesday, August 18, 2026, at 2:10 AM PDT , Flexera identified an issue affecting the Flexera One APAC Production environment that prevented customers from accessing Flexera One services. During the impact window, customers experienced login failures, session refresh failures, and issues retrieving user-related application data, resulting in widespread service access unavailability across the platform.\n\nInitial investigation indicated that the impact was limited to IT Asset Management. Further analysis determined that the issue involved a shared platform service supporting authentication and authorization across Flexera One, resulting in a broader impact across the APAC environment.\n\nOur technical teams identified an incorrect infrastructure scheduling configuration that prevented a critical platform service from operating as intended. The configuration was corrected, restoring the affected services. Following remediation, health checks returned successfully, application login and navigation functionality recovered, and teams completed post-restoration validation.\n\nThe environment was subsequently confirmed healthy and placed under continued monitoring to ensure ongoing stability.\n\n\u200c\n\n**Root Cause**\n\n\u200c\n\nThe incident was caused by an incorrect infrastructure scheduling configuration within a critical shared identity and access service. The configuration prevented the service from being scheduled onto the intended infrastructure, resulting in service unavailability and failures affecting authentication and authorization requests.\n\n\u200c\n\n**Remediation Actions**\n\n\u200c\n\nOur teams performed the below actions to restore the services to normalcy:\n\n* Configuration Correction: Corrected the infrastructure scheduling configuration affecting the impacted platform service.\n\n* Service Restoration: Restored the affected platform services and confirmed successful health checks.\n\n* Post-Restoration Validation: Performed automated health checks and platform validation to confirm application login, navigation, and related functionality had recovered.\n\n* Continued Monitoring: The environment was monitored following restoration to confirm ongoing platform stability.\n\n\u200c\n\n**Future Preventative Measures**\n\n\u200c\n\n* Pre-Deployment Validation: Implement additional validation checks for infrastructure scheduling and placement configurations.\n\n* Deployment Safeguards and Resiliency Controls: Review and strengthen deployment safeguards, configuration governance, and resiliency controls to prevent configuration mismatches from impacting service availability.\n\n* Expanded Post-Change Verification: Review and enhance the post-deployment and post-maintenance validation processes to ensure critical platform services are functioning correctly following infrastructure changes.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "critical",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-08-18T05:13:38.635-07:00",
"resolved_inferred": false,
"started_at": "2026-08-18T02:32:31.000-07:00",
"state": "postmortem",
"title": "Flexera One - APAC -  Access discruption",
"updated_at": "2026-08-31T04:35:36.953-07:00",
"url": "https://stspg.io/9zxc8zlfb5c6"
},
{
"body": "**Description:** Flexera One IT Asset Management \u2013 EU \u2013 Reconciliation Failures\n\n**Timeframe:**  August 12, 2026, 5:50 PM PDT to August 12, 2026, 8:29 PM PDT\n\n\u200c\n\n**Incident Summary**\n\n\u200c\n\nOn Wednesday, August 12, 2026, at 5:50 PM PDT , Flexera identified an issue affecting reconciliation process within the Flexera One IT Asset Management \\(ITAM\\) EU Production environment. During the impact period, Some reconciliation jobs experienced failures when they were assigned to a specific processing instance that had not completed its required configuration.\n\nInitial investigation determined that the issue was isolated to a single server that did not complete its configuration successfully. As a result, the affected server was unable to establish communication with a required dependent service. Reconciliation jobs processed by this server failed, while jobs routed to other correctly configured processing servers continued to complete successfully.\n\nTechnical teams investigated the issue, identified the configuration discrepancy, completed the required connectivity configuration, and validated successful communication with the dependent service at 8:29 PM PDT. Following remediation, reconciliation processing resumed normally, affected workloads were reprocessed, and continued monitoring confirmed service stability. After monitoring the services to ensure stability , the issue was declared as restored at 10:56 PM PDT.\n\n\u200c\n\n**Root Cause**\n\n\u200c\n\nA batch processing server entered service with an incomplete configuration due to an initialization process that did not fully complete. As part of the required deployment sequence, network connectivity configuration must be established before the IT Asset Management processing components become available to process workloads.\n\nBecause the required connectivity configuration was not completed on the affected server, it was unable to communicate with a dependent reconciliation service. When reconciliation jobs were assigned to this server, processing failed, resulting in reconciliation failures for workloads routed through the misconfigured instance.\n\nContributing Factors\n\n* The issue was isolated to a single batch processing server that did not complete its full initialization and configuration sequence successfully.\n* The server became available to process workloads before all required connectivity configuration steps had been fully validated.\n* Required service connectivity configuration was not fully established prior to the server entering normal processing operations.\n* Impact was limited to reconciliation jobs routed through the affected server, while jobs processed by correctly configured servers continued to complete successfully.\n* Existing deployment validation controls did not detect the incomplete configuration state before the server began processing customer workloads.\n\n\u200c\n\n**Remediation Actions**\n\n\u200c\n\n* Initialization Process Rerun: The failed initialization process was rerun successfully on the affected batch processing instance.\n* Configuration Completion: All required post-installation configuration steps were completed, restoring the instance to the expected configuration.\n* Connectivity Restoration: Connectivity between the IT Asset Management processing environment and the reconcile service was restored and validated.\n* Service Validation: Teams confirmed that reconciliation failures had stopped occurring and continued monitoring the environment to verify stable operation.\n\n\u200c\n\n**Future Preventative Measures**\n\n\u200c\n\n* Deployment Sequence Improvements - Update deployment procedures to ensure processing services remain unavailable until all initialization and configuration activities have completed successfully.\n* Post-Deployment Validation Enhancements - Implement additional validation checks to verify that all required configuration steps have completed before services enter production operation.\n* Deployment Control Review - Review and strengthen deployment controls to ensure service startup sequencing and configuration requirements are consistently enforced across future releases.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-08-12T22:57:52.297-07:00",
"resolved_inferred": false,
"started_at": "2026-08-12T18:08:56.642-07:00",
"state": "postmortem",
"title": "Flexera One \u2013 IT Asset Management \u2013 EU \u2013 Reconciliation Failures",
"updated_at": "2026-08-26T23:51:27.932-07:00",
"url": "https://stspg.io/lgy2978cdh3p"
},
{
"body": "**Description:** Flexera One - Cloud Cost Optimization \\(CCO\\) - North America - Service Degradation Affecting Certain Functionality\n\n**Timeframe:** August 11, 2026, 9:51 AM PDT \u2013 August 11, 2026, 10:45 AM PDT\n\n**Incident Summary**\n\nOn August 11, 2026, customers may have experienced issues accessing certain Cloud Cost Optimization \\(CCO\\) functionality in the North America region. Impacted pages could remain in a continuous loading state and fail to render correctly.\n\nTechnical teams began investigating after identifying that certain application requests were not completing successfully. The impact was limited to specific functionality rather than the entire CCO service. \n\nDuring the investigation, technical teams determined that dependent application services were unable to successfully retrieve required configuration information from a supporting service component, resulting in rendering issues for affected functionality.\n\nCorrective configuration changes were implemented and updated service components were deployed. Following deployment, technical teams validated recovery and continued monitoring the environment to confirm stable operation. All affected functionality subsequently returned to normal operation.\n\n**Root Cause**\n\nDuring a recent service migration, a supporting application component was deployed with an incorrect resource configuration. As processing demand increased, the affected service components exhausted available resources and became unavailable.\n\nAs a result, dependent application services were unable to successfully retrieve required configuration information needed to process certain requests. This resulted in rendering issues affecting specific Cloud Cost Optimization \\(CCO\\) functionality.\n\n**Remediation Actions**\n\nThe following actions were taken during the incident response:\n\n1. Incident Investigation Initiated: Technical teams investigated reports of affected CCO functionality remaining in a continuous loading state and failing to render correctly.\n2. Service Dependency Analysis Performed: Technical teams identified that dependent application services were unable to successfully retrieve required configuration information from a supporting service component.\n3. Configuration Issue Identified: Investigation determined that a supporting application component had been deployed with an incorrect resource configuration following a recent service migration.\n4. Corrective Configuration Changes Applied: Technical teams implemented configuration changes to restore the intended operating parameters for the affected service components.\n5. Service Recovery Validation Performed: Updated service components were deployed, affected functionality was validated, and the environment was monitored to confirm stable operation before incident closure.\n\n**Future Preventative Measures**\n\nFollowing this incident, technical teams reviewed the migration and deployment process associated with the affected service component.\n\n1. Configuration Validation Enhancements: An additional validation step has been added to verify that service configurations are appropriate for forecasted processing requirements following migration activities. This additional review is intended to help identify configuration discrepancies before they can affect production workloads.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-08-11T13:26:38.231-07:00",
"resolved_inferred": false,
"started_at": "2026-08-11T10:26:42.000-07:00",
"state": "postmortem",
"title": "Flexera One - Cloud Cost Optimization (CCO) - NAM - Service Degradation",
"updated_at": "2026-08-14T13:03:21.831-07:00",
"url": "https://stspg.io/dk3gp7fr5l3d"
},
{
"body": "**Description:** Flexera Support Cases \u2013 Delayed or Failed Email Notifications\n\n**Timeframe:**  August 5, 2026, 12:00 AM PDT to August 13, 2026, 11:10 AM PDT\n\n\u200c\n\n**Incident Summary**\n\n\u200c\n\nOn August 5, 2026, our teams identified an issue affecting email notifications for Flexera Support cases. During the incident, some customers did not receive email notifications when updates were made to their support cases, which may have delayed awareness of case activity and progress.\n\nOur technical teams investigated the issue and determined that an issue with the email delivery configuration was causing some outbound support case notifications to be rejected by recipient email systems.\n\nAs an interim measure, customers were advised to access the Flexera Community Support Case portal directly to review their case status and updates. Our teams implemented a permanent change to the email delivery configuration and subsequently validated notification delivery.\n\nFollowing the remediation, email notifications were confirmed to be operating as expected, and the incident was resolved on August 13, 2026, at 11:10 AM PDT.\n\n\u200c\n\n**Root Cause**\n\n\u200c\n\n* Email Delivery Configuration: An issue with the configuration used to deliver customer-facing support case notifications caused some emails to be rejected by recipient email systems.\n* Notification Delivery: The affected configuration resulted in some customers not receiving email notifications when updates were made to their support cases.\n* Customer Impact: The issue affected email notifications only; customers could continue to access their support cases directly through the Flexera Community Support Case portal.\n\n\u200c\n\n**Remediation Actions**\n\n\u200c\n\n* Updated Email Delivery Configuration: The email delivery configuration was permanently  updated to use the appropriate primary email delivery infrastructure.\n* Notification Validation: Teams validated outbound notification delivery following the configuration change.\n* Customer Workaround: During the incident, customers were advised to access the Flexera Community Support Case portal directly to review case status and updates.\n* Service Monitoring: Teams monitored email notification delivery following remediation to confirm continued stability.\n\n\u200c\n\n**Future Preventative Measures**\n\n* Maintain Approved Email Routing: Customer-facing support case notifications will continue to use the primary email delivery infrastructure established through the permanent remediation.\n* Email Delivery Monitoring: Enhance monitoring and alerting for outbound support case notification failures and rejected messages to provide earlier detection of potential delivery issues.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-08-13T22:41:52.974-07:00",
"resolved_inferred": false,
"started_at": "2026-08-06T21:55:10.252-07:00",
"state": "postmortem",
"title": "Flexera Support Cases - Delayed or Failed Email Notifications for Case Updates",
"updated_at": "2026-08-21T02:14:49.542-07:00",
"url": "https://stspg.io/gc2f663xhz0z"
},
{
"body": "**Description:** Flexera One - IT Asset management - APAC - Data loading failures/slowness\n\n**Timeframe:**  August 4, 2026, 4:04 AM PDT to August 4, 2026, 4:25 AM PDT\n\n\u200c\n\n**Incident Summary**\n\n\u200c\n\nOn Tuesday, August 4, 2026, our teams identified an issue affecting a subset of customers in the APAC Production environment. During the incident, some customers experienced intermittent application slowness, temporary unresponsiveness, and failures when loading data within Flexera One IT Asset Management.\n\nOur technical teams investigated the issue and determined that an unexpected database resource exhaustion event affected the production environment. The condition resulted in temporary performance degradation for a subset of customer tenants.\n\nThe database condition resolved automatically after approximately 20 minutes, and application performance returned to normal without requiring direct engineering intervention. Our teams continued to monitor the environment following recovery and confirmed that affected services were operating normally with no ongoing customer impact..\n\n\u200c\n\n**Root Cause**\n\n\u200c\n\n* Unexpected Database Resource Exhaustion: A temporary increase in database resource consumption exhausted the available temporary database storage capacity used by the production database cluster.\n* High Resource-Consuming Queries: It was identified that specific unexpected query run manually by a user that consumed significant amounts of temporary database resources, resulting in a temporary resource exhaustion condition.\n* Application Performance Impact: The resource exhaustion affected database operations and resulted in intermittent application slowness, unresponsiveness, and data loading failures for a subset of customers.\n\n\u200c\n\n**Remediation Actions**\n\n\u200c\n\n* Automatic Service Recovery: The database resource condition cleared automatically, allowing application services to return to normal operation.\n* Service Monitoring: Teams monitored the environment following recovery and confirmed that services remained stable.\n* Incident Analysis: Engineering teams reviewed the database activity and resource utilization associated with the event to identify the contributing queries and understand the behaviour. \n\n\u200c\n\n**Future Preventative Measures**\n\n\u200c\n\n* Resource Governance: Evaluate additional SQL Server resource governance controls to reduce the potential impact of high-resource database queries.\n* Alert Response Improvements: Review alert handling procedures to improve detection, assessment, and response to similar database resource conditions.\n* Operational Improvements: The query has been reviewed by our technical teams and measure have been put in place to avoid a similar scenario in the future.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-08-05T01:28:39.811-07:00",
"resolved_inferred": false,
"started_at": "2026-08-05T01:28:39.751-07:00",
"state": "postmortem",
"title": "Flexera One - IT Asset management -  APAC - Data loading failures/slowness",
"updated_at": "2026-08-21T02:05:10.543-07:00",
"url": "https://stspg.io/z0fwqsdk6mtg"
},
{
"body": "**Description:** Flexera One- IT Asset management - EU - Inventory upload errors\n\n**Timeframe:**  July 31, 2026, 6:00 PM PDT to August 4, 2026, 4:09 AM PDT\n\n\u200c\n\n**Incident Summary**\n\n\u200c\n\nOn Sunday, August 2, 2026, at 11:42 PM PDT, Flexera identified an issue affecting inventory package \\(\\*.zip\\) uploads and select beacon communications within the Flexera One IT Asset Management EU Production environment. Affected customers experienced failed inventory package uploads, including HTTP 504 \\(Gateway Timeout\\) and HTTP 405 \\(Method Not Allowed\\) responses. The issue also disrupted select beacon-to-platform communications, causing delays or failures in inventory processing and inventory data ingestion.\n\nThe investigation determined that earliest impact started at 6:00 PM PDT on July 31st , 2026. Analysis identified a communication issue within a critical service routing path that prevented a subset of requests from successfully reaching downstream processing services. This resulted in repeated retries, delayed inventory processing, and backlog accumulation.\n\nTo mitigate the issue and restore service, traffic was redirected to an alternate communication path. Following this change, beacon communications stabilized, inventory uploads resumed processing successfully, and the inventory backlog began to decrease. By 3:42 AM PDT on August 3, 2026, inventory uploads were completing successfully and no new errors were observed.\n\nThe environment remained under extended monitoring while processing volumes returned to normal operating levels. At 4:09 AM PDT on August 4, 2026, Flexera confirmed that beacon communications, inventory processing, and associated platform services were operating normally, the backlog had returned to expected thresholds, and the incident was formally declared resolved.\n\n\u200c\n\n**Root Cause**\n\n\u200c\n\nOn 31 July 2026, a certificate update to a critical service communication path created a configuration mismatch between platform components. The issue surfaced later, when application components were restarted and began establishing new connections using the updated configuration. This caused handshake errors and communication failures, resulting in service unavailability. Consequently, requests from customer beacons to upload inventory package \\(\\*.zip\\) files to the EU Production environment could not be processed, causing upload failures and preventing inventory data from being received and processed. A subset of other beacon-to-platform interactions was also affected, although customer impact was not observed immediately.\n\nService was restored by redirecting traffic to an alternate communication path that was not affected by the certificate-related configuration issue. Inventory uploads and beacon communications then resumed normal operation.\n\nContributing Factors\n\n* Intermittent Request Failures: Failed connection establishment attempts caused a portion of requests to be unsuccessful before reaching downstream processing services.\n\n* Inventory Processing Delays: Communication failures prevented some inventory packages and associated imports from processing successfully, resulting in delayed inventory updates.\n\n* Retry-Driven Backlog Growth: Automated retry mechanisms enabled many transactions to eventually succeed but contributed to increased processing delays and backlog accumulation while the issue remained active.\n\n\u200c\n\n**Remediation Actions**\n\n\u200c\n\nTo restore service, technical teams performed below actions:\n\n* Traffic Re-routing: Redirected traffic away from the affected service path to restore stable beacon communications and inventory package uploads.\n\n* Platform Monitoring: Performed continuous monitoring of beacon communications, inventory uploads, and processing activity to verify service recovery.\n\n* Backlog Recovery Management: Monitored inventory processing throughput and validated that processing backlogs steadily reduced and returned to expected operating thresholds.\n\n* Stability Validation: Conducted extended monitoring following restoration to confirm sustained platform stability and successful inventory processing.\n\n* Service Restoration Confirmation: At 4:09 AM PDT on August 4, 2026, confirmed that beacon communications, inventory processing, and associated services were operating normally.\n\n\u200c\n\n**Future Preventative Measures**\n\n\u200c\n\n* Service Communication Resilience: Review and enhance the affected communication path to improve reliability, resiliency, and recovery from future connectivity-related disruptions.\n\n* Strengthened Change Validation: Enhance validation and testing processes for infrastructure, configuration, certificate, and service-path changes to identify communication mismatches before deployment into production environments.\n\n* Improved Detection and Alerting: Enhance monitoring and alerting capabilities to provide earlier identification of sustained communication disruptions, elevated retry activity, abnormal error-rate patterns, and emerging processing backlogs. These improvements will enable faster detection, investigation, and remediation of issues before they develop into broader customer impact.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-08-04T04:09:12.946-07:00",
"resolved_inferred": false,
"started_at": "2026-08-02T23:56:20.688-07:00",
"state": "postmortem",
"title": "Flexera One- IT Asset management - EU  - Inventory upload and Beacon communication issues",
"updated_at": "2026-08-18T05:43:21.118-07:00",
"url": "https://stspg.io/bqxvlf3421k1"
},
{
"body": "**Description:** Software Vulnerability Research \\(SVR\\) - Service Disruption\n\n**Timeframe:**  July 28, 2026, 2:00 PM PDT to July 29, 2026, 2:53 AM PDT\n\n\u200c\n\n**Incident Summary**\n\nOn Tuesday, 28 July 2026, at 2:00 PM PDT, the Flexera Software Vulnerability Research\\(SVR\\) production environment experienced a service disruption that affected API availability and application functionality. The incident occurred when the primary database instance supporting the SVR platform became unavailable. Existing database connections were unexpectedly terminated, and application servers were unable to establish new connections. As a result, customers were unable to reliably access SVR services and APIs during the incident.\n\nDuring the investigation, our teams determined that the disruption was caused by an unplanned failover initiated by the cloud service provider. According to the provider\u2019s event history, an infrastructure issue was detected on the primary database host, prompting the provider to automatically initiate the failover process.\n\nOur technical teams immediately engaged the cloud service provider to investigate the incident and restore service. Following recovery and validation activities, all production services were successfully restored and returned to normal operation by 2:53 AM PDT on 29 July 2026.\n\n\u200c\n\n**Root Cause**\n\n\u200c\n\nThe incident was triggered by an unplanned failover of the production database, initiated by the cloud service provider in response to an infrastructure-level issue affecting the primary database host.\n\nDuring the failover, the primary database instance became temporarily unavailable. This caused active database connections to terminate abruptly, preventing application services from establishing new connections. As a result, API requests failed and service availability across the SVR platform was impacted.\n\nFlexera has formally requested a detailed Root Cause Analysis \\(RCA\\) from the cloud service provider to determine the precise infrastructure condition that triggered the failover and to identify opportunities to prevent recurrence.\n\n\u200c\n\n**Contributing Factors**\n\n\u200c\n\nDuring the investigation, the following factors were identified as potentially contributing to the overall impact of the incident:\n\n* A significantly elevated volume of client connection attempts was observed originating from customer environments during the incident period.\n\n* Connection volumes exceeded normal operating levels, increasing load on backend services while database services were recovering.\n\n* Elevated connection activity may have amplified the impact of the database failover and increased recovery complexity.\n\n* Existing application retry and connection behaviours generated additional connection demand during database recovery activities.\n\n\u200c\n\n**Remediation Actions**\n\n\u200c\n\n* Cloud provider engagement: Our teams immediately engaged the cloud service provider and worked directly with their on-call engineers throughout the recovery effort.\n\n* Database recovery: Executed a controlled database failover/reboot in coordination with the cloud service provider.\n\n* Restored connectivity: Re-established database availability and connectivity to the production environment.\n\n* Application recovery validation: Confirmed that application services successfully reconnected to backend database services.\n\n* Service validation: Verified recovery of APIs and critical SVR application functionality.\n\n* Post-recovery monitoring: Implemented enhanced monitoring of database health, application performance, and connection activity following restoration.\n\n\u200c\n\n**Future Preventative Measures**\n\n\u200c\n\n* Continue monitoring customer connection volumes and backend database connection utilisation.\n\n* Review and optimise database connection pooling, connection management, retry logic, and failover handling to improve resilience during future infrastructure events.\n\n* Complete the cloud service provider support engagement and review the provider's detailed Root Cause Analysis once available.\n\n* Conduct an internal post-incident review and implement additional preventive measures identified through the review process.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "critical",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-29T02:52:54.749-07:00",
"resolved_inferred": false,
"started_at": "2026-07-28T21:36:06.142-07:00",
"state": "postmortem",
"title": "Software Vulnerability Research (SVR) - Service Disruption",
"updated_at": "2026-08-11T03:32:17.731-07:00",
"url": "https://stspg.io/sy76mhkrh4t3"
},
{
"body": "**Description:** Flexera One - APAC - Intermittent Service Disruption\n\n## Timeframe\n\n**Timeframe:** July 28, 2026, 3:16 PM PDT - July 28, 2026, 5:39 PM PDT\n\n## Incident Summary\n\nOn Tuesday, July 28, 2026, at 3:49 PM PDT, customers using Flexera One services in the APAC region began experiencing intermittent service disruptions affecting multiple applications and platform capabilities. Customers may have encountered login issues, application errors, failed page loads, API errors, and intermittent access to certain Flexera One functionality.\n\nDuring the incident, multiple services experienced intermittent communication failures with shared platform services. As a result, customers experienced inconsistent application behavior, with some requests succeeding while others failed.\n\nTechnical teams investigated the issue across affected applications and platform components to determine the scope and source of the failures. The investigation identified a connectivity issue affecting communication between application services and a shared platform dependency. Corrective configuration changes were implemented to restore service communication and stabilize affected functionality.\n\nFollowing implementation of the corrective changes, technical teams validated functionality across impacted applications and confirmed that services had returned to normal operation. Continued monitoring showed stable service behavior, and the incident was resolved at 5:40 PM PDT on July 28, 2026.\n\n## Root Cause\n\nInvestigation determined that the incident was caused by a connectivity issue introduced during a platform infrastructure migration.\n\nFollowing the migration, certain application services continued using legacy connection ports when communicating with shared platform services. When affected services restarted, they attempted to communicate using endpoint and port combinations that were no longer aligned with the updated platform configuration.\n\nThis configuration mismatch resulted in intermittent communication failures between application services and shared platform components, causing customer-facing application errors and service disruptions across multiple Flexera One services in the APAC region.\n\n## Remediation Actions\n\nThe following actions were taken during the incident response:\n\n1. **Incident Investigation Initiated:** Technical teams investigated reports of intermittent failures affecting multiple Flexera One services in the APAC region and assessed the scope of customer impact.\n2. **Dependency Analysis Performed:** Technical teams reviewed communication paths between impacted applications and shared platform services to identify the source of the connectivity failures.\n3. **Configuration Mismatch Identified:** Investigation determined that certain services were attempting to communicate through legacy ports that were not aligned with the updated platform configuration following the migration.\n4. **Platform Configuration Updated:** The required ports were added to the updated platform configuration, restoring connectivity for the affected services.\n5. **Service Recovery Validation:** Technical teams validated functionality across impacted applications and confirmed that service communication and customer-facing functionality had returned to normal operation.\n\n## Future Preventative Measures\n\nThe following improvements have been identified to further reduce the risk of similar incidents in the future:\n\n1. **Platform Configuration Alignment:** Platform connectivity configurations were updated to ensure application services can communicate with shared platform components using the required connection ports. Maintaining alignment between service connectivity requirements and platform configurations helps reduce the risk of similar connectivity issues following future infrastructure changes.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-28T19:05:31.457-07:00",
"resolved_inferred": false,
"started_at": "2026-07-28T16:14:54.944-07:00",
"state": "postmortem",
"title": "Flexera One - APAC - Service Disruption",
"updated_at": "2026-08-12T16:01:49.434-07:00",
"url": "https://stspg.io/wfwr9zq8npq9"
},
{
"body": "**Description:** Snow Atlas - APAC - Login Failures and 404 Errors\n\n**Timeframe:**  July 26, 2026, 5:00 PM PDT to July 26, 2026, 6:15 PM PDT\n\n\u200c\n\n**Incident Summary**\n\n\u200c\n\nOn Sunday, 26 July 2026, Flexera experienced a service disruption affecting Snow Atlas customers in the APAC Production environment, resulting in login failures and temporary inability to access Agreement pages. The issue was isolated to the APAc region, with no impact to other production regions.\n\nInvestigation by our technical teams determined that the disruption was caused by a backend service communication issue that prevented Agreement data from being retrieved successfully. The condition was consistent with a runtime synchronization issue following recent platform infrastructure maintenance.\n\nFlexera teams validated platform health, restored the affected services, and confirmed successful recovery. Customer access was fully restored, and post-recovery monitoring confirmed stable operations with no recurrence of the underlying service communication errors.\n\nA detailed post-incident review was completed, and corrective measures have been implemented to further strengthen platform resilience and recovery processes.\n\n\u200c\n\n**Root Cause**\n\n\u200c\n\nThe incident was caused by an application communication failure within the APAC Production environment. A component responsible for processing requests between internal application services did not successfully handle Agreement-related requests, preventing those requests from being completed and resulting in customer-facing 404 errors and failures accessing Agreement pages.\n\n\u200c\n\n**Contributing Factors**\n\n\u200c\n\n* The application communication failure occurred following a recent infrastructure upgrade to the APAC Production environment.\n* Although the application services and underlying infrastructure remained healthy, an internal service responsible for processing Agreement requests did not fully recover after the upgrade.\n* Existing health checks validated application and infrastructure availability but did not verify the successful initialization of the internal communication path following the upgrade.\n* The issue was therefore not detected until customers began experiencing login failures and HTTP 404 errors when accessing Agreement pages.\n\n\u200c\n\n**Remediation Actions**\n\n\u200c\n\nTo restore service, technical teams:\n\n* Confirmed the underlying infrastructure remained healthy throughout the incident.\n* Identified the failed internal service communication affecting Agreement requests.\n* Restarted the affected application services.\n* Validated successful recovery across impacted tenants.\n* Continued monitoring to confirm service stability.\n\n\u200c\n\n**Future Preventative Measures**\n\n\u200c\n\n* Enhanced Post-Upgrade Validation: Introduce additional validation checks following infrastructure upgrades to verify that critical application components and internal service communications are functioning as expected. \n* Improved Monitoring and Alerting: Enhance monitoring and alerting to detect internal service communication failures and responder registration issues more quickly. \n* Synthetic Health Checks: Implement synthetic health checks that validate critical customer workflows, including Agreement page functionality, following infrastructure changes.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "critical",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-26T19:00:22.790-07:00",
"resolved_inferred": false,
"started_at": "2026-07-26T17:11:18.717-07:00",
"state": "postmortem",
"title": "Snow Atlas - APAC - Login Failures and 404 Errors",
"updated_at": "2026-08-07T01:47:00.838-07:00",
"url": "https://stspg.io/459ynts50k3j"
},
{
"body": "**Description:** Flexera One \u2013 IT Visibility \u2013 EU \u2013 Failure to Load Reports\n\n**Timeframe:** July 22, 2026, 3:11 AM PDT \u2013 July 22, 2026, 11:04 PM PDT  \n\n**Incident Summary**\n\nOn July 22, 2026, at approximately 3:11 AM PDT, an issue began affecting IT Visibility reporting functionality in the Flexera One EU production environment. During the affected period, customers in the EU region experienced failures when attempting to load IT Visibility reports and dashboards, including Software Inventory, Hardware Inventory, Technology Intelligence \\(TI\\), Sustainability, and FinOps reporting views.\n\nTechnical teams began investigating after multiple customer reports were received regarding report loading failures across the EU region. Initial analysis confirmed the issue was limited to the EU production environment, while equivalent reporting functionality in other regions continued to operate normally.\n\nThe investigation determined that an authentication-related issue had occurred within the EU reporting environment. As a result, affected reports and dashboards were unable to retrieve the data required to load successfully, causing report loading failures for impacted customers.\n\nTechnical teams implemented corrective updates within the affected reporting environment and performed validation activities across impacted customer organizations. Following these actions, reports and dashboards resumed normal operation and customers were once again able to access reporting data successfully.\n\nFollowing validation of affected customer environments and confirmation that reports were loading successfully, service functionality was restored and the incident was considered resolved. Technical teams continued monitoring following restoration to ensure ongoing service stability.\n\n**Root Cause**\n\nThe incident was caused by an authentication-related issue within the EU reporting environment. This issue prevented affected reports and dashboards from successfully retrieving the data required to display results, resulting in report loading failures for impacted customers.\n\n**Remediation Actions**\n\nThe following actions were taken during the incident response:\n\n1. Regional Impact Assessment: Technical teams investigated customer reports and confirmed the issue was isolated to the EU production environment. Additional validation was performed across other regions to verify that reporting functionality remained operational outside the affected environment.\n2. Authentication Failure Analysis: Technical teams reviewed report processing, authentication workflows, reporting workspaces, and recent platform changes to identify the source of the failures. The investigation determined that affected reports were unable to successfully authenticate when retrieving underlying data required for report rendering.\n3. Restoration and Validation Activities: Technical teams applied updates across affected customer environments and reporting workspaces to restore successful report data retrieval. Following implementation, impacted customer organizations were validated to confirm reports and dashboards were loading successfully and displaying data as expected.\n4. Post-Recovery Verification: Following restoration, additional monitoring and validation activities were conducted to verify continued report functionality, confirm successful recovery, and ensure service stability across the EU production environment.\n\n**Future Preventative Measures**\n\nThe following follow-up actions were identified during the incident:\n\n\u2022\tAuthentication Update Process Review: Technical teams are reviewing the authentication update process to understand why the update was not successfully applied across all affected reporting workspaces in the EU production environment and to identify improvements that will help prevent similar issues in future updates.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-23T00:52:01.928-07:00",
"resolved_inferred": false,
"started_at": "2026-07-22T03:33:52.886-07:00",
"state": "postmortem",
"title": "Flexera One- IT Visibility- EU - Failure to load reports",
"updated_at": "2026-08-03T11:40:13.157-07:00",
"url": "https://stspg.io/s2tkp8s858rh"
},
{
"body": "**Description:** Flexera One \u2013 IT Asset Management \u2013 North America \u2013 Degraded Performance\n\n**Timeframe:** July 20, 2026, 8:13 PM PDT \u2013 July 22, 2026, 8:41 AM PDT\n\n\u200c\n\n**Incident Summary**\n\nOn July 20, 2026, at approximately 8:13 PM PDT, an issue began affecting Flexera One IT Asset Management customers in the North America production environment following completion of the production release.\n\nDuring the affected period, some customers experienced intermittent application slowness, delayed processing, reconcile failures, timeout conditions, and degraded performance across inventory-related workflows. Technical teams investigated performance and reliability issues observed following the release and worked to determine the underlying cause and scope of impact.\n\nThe investigation determined that the observed behavior was associated with a defect previously identified in the release. Corrective changes intended to address the defect were included as part of the release activity; however, the corrective changes did not successfully apply across all North America production databases during the deployment process. As a result, some environments continued to experience significantly increased database processing activity during inventory-related operations, contributing to elevated resource consumption, processing delays, timeout conditions, and degraded application performance.\n\nTechnical teams worked to successfully deploy the corrective changes across the affected North America production databases. Following deployment, teams performed validation activities and monitored processing workflows to confirm recovery.\n\nBy July 22, 2026, at approximately 8:41 AM PDT, validation confirmed successful processing of previously impacted workloads and that the previously observed failure patterns were no longer occurring. The incident was considered resolved and technical teams continued monitoring to confirm service stability.\n\n\u200c\n\n**Root Cause**\n\nThe incident was caused by corrective changes intended to address a previously identified defect not successfully applying across all North America production databases during the release deployment process.\n\nAs a result, the underlying defect remained active in affected environments and caused significantly more database processing activity than intended during inventory-related processing operations. This increased workload resulted in elevated resource consumption, processing delays, timeout conditions, and degraded platform performance.\n\nThese conditions contributed to intermittent application slowness, delayed processing, reconcile failures, and degraded performance affecting some IT Asset Management functionality within the North America production environment.\n\n\u200c\n\n**Remediation Actions**\n\n1. **Incident Investigation Initiated:** Technical teams investigated reports of application slowness, processing delays, reconcile failures, and degraded platform performance affecting the North America production environment.\n2. **Impact Assessment Performed:** Teams reviewed customer-reported symptoms, system performance data, processing activity, and platform behavior to determine the scope and nature of the issue.\n3. **Cause Identified:** Investigation determined that corrective changes intended to address a previously identified defect did not successfully apply across all North America production databases during the release deployment process.\n4. **Processing Behavior Corrected:** Technical teams deployed corrections designed to eliminate the inefficient processing behavior that was contributing to elevated database workload, processing delays, timeout conditions, and degraded application performance.\n5. **Recovery Validated:** Teams validated the effectiveness of the deployed changes through monitoring and successful execution of previously impacted processing activities.\n6. **Post-Restoration Monitoring Performed:** Additional monitoring confirmed that the previously observed failure patterns and performance degradation were no longer occurring and that service stability had been restored.\n\n\u200c\n\n**Future Preventative Measures**\n\nBased on the investigation, the following follow-up activities have been identified:\n\n* **Processing Logic Improvements:** Technical teams implemented and validated corrective changes to the inventory-processing logic responsible for the increased database workload observed following the release. These changes were designed to eliminate the inefficient processing behavior that contributed to processing delays, reconcile failures, timeout conditions, and degraded application performance.\n* **Critical Hotfix Deployment Process Review:** Technical teams will review the deployment process for critical database hotfixes and evaluate alternative approaches to help ensure required corrective changes are successfully applied during future release activities.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-22T09:16:50.127-07:00",
"resolved_inferred": false,
"started_at": "2026-07-22T01:25:49.902-07:00",
"state": "postmortem",
"title": "Flexera One- IT Asset management- US - Degraded performance",
"updated_at": "2026-08-04T13:57:39.110-07:00",
"url": "https://stspg.io/n9wdjbhwd6d6"
},
{
"body": "**Description:** Flexera Spot \u2013 AWS Billing Service Disruption\n\nTimeframe: July 17, 2026, 1:33 AM PDT \u2013 July 18, 2026, 6:57 AM PDT  \n\n**Incident Summary**\n\nOn July 17, 2026, at approximately 1:33 AM PDT, Flexera technical teams became aware of an AWS issue affecting estimated billing and usage data displayed within AWS Billing and Cost Management, Cost Explorer, and Cost and Usage Reports.\n\nGiven the nature of the AWS service disruption, technical teams initiated an investigation to determine whether Spot billing, cost analysis, or savings calculations for AWS customers were affected. During the investigation, teams reviewed the billing and pricing data used by Spot, validated relevant cost and pricing tables, monitored AWS updates, and assessed whether any customer-facing billing or savings data had been impacted.\n\nTechnical teams validated the billing and pricing data used by Spot and confirmed that the relevant cost and pricing tables were operating as expected. The investigation did not identify any issues within the data used by Spot, and no incorrect billing or savings calculations were observed.\n\nAWS subsequently confirmed that the issue originated within AWS Billing and Cost Management services, identified and mitigated the underlying issue, and completed data recovery and backfill activities.\n\nBased on the validation performed throughout the investigation, no customer impact was observed within Spot. Following confirmation of AWS recovery and completion of post-recovery validation activities, the incident was considered resolved.  \n\n**Root Cause**\n\nThe incident was caused by an AWS service issue affecting estimated billing and usage data displayed within AWS Billing and Cost Management, Cost Explorer, and Cost and Usage Reports. AWS identified and mitigated the underlying issue and completed data recovery activities.  \n\n**Remediation Actions**\n\nThe following actions were taken during the incident response:  \n\u2022\tIncident Investigation Initiated: Technical teams began investigating the AWS billing and usage data issue to determine whether Spot customers were affected.  \n\u2022\tImpact Assessment Performed: Technical teams evaluated whether the AWS service disruption affected Spot billing analysis and savings calculations.  \n\u2022\tData Validation Completed: Teams reviewed and validated the billing and pricing data used by Spot and confirmed that relevant cost and pricing tables were operating as expected.  \n\u2022\tAWS Recovery Monitored: Technical teams monitored AWS communications, service updates, and recovery activities throughout the incident.  \n\u2022\tPost-Recovery Validation Performed: Additional validation was completed following AWS recovery and data backfill activities to confirm the integrity of billing and pricing data used by Spot.  \n\u2022\tIncident Resolved: The incident was closed after AWS confirmed recovery and technical validation identified no impact within Spot.   \n\n**Future Preventative Measures**\n\nThe following follow-up actions were identified during the incident:\n\n\u2022\tAWS Collaboration and Review: Continue working closely with AWS to understand the underlying cause of provider-managed incidents and review any recommendations or learnings resulting from their investigation.  \n\u2022\tResiliency and Validation Review: Review opportunities to further strengthen validation processes for billing-related data to support faster impact assessment and response during similar third-party service disruptions in the future.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-19T05:59:04.809-07:00",
"resolved_inferred": false,
"started_at": "2026-07-17T03:24:31.000-07:00",
"state": "postmortem",
"title": "Flexera- Spot- All regions - Incorrect billing data for AWS customers",
"updated_at": "2026-07-30T11:11:23.034-07:00",
"url": "https://stspg.io/rgkjf2ypvr0n"
},
{
"body": "**Description:** Flexera One \u2013 IT Asset Management \u2013 APAC & EU \u2013 Slowness and Degraded Performance\n\n**Timeframe:** July 13, 2026, 1:00 PM PDT \u2013 July 18, 2026, 7:10 AM PDT  \n\n**Incident Summary**\n\nOn July 13, 2026, at approximately 1:00 PM PDT, an issue began affecting Flexera One IT Asset Management customers in the APAC and Europe production environments following deployment of the 2026 R1.1 release.\n\nDuring the affected period, some customers experienced intermittent application slowness, delayed processing, reconcile failures, and degraded performance when accessing certain areas of the application or performing inventory-related activities. As customer reports increased, technical teams began investigating common patterns across the reported symptoms to determine the underlying cause.\n\nThe investigation identified a defect introduced as part of the Flexera One IT Asset Management 2026 R1.1 release. Under certain conditions, the defect resulted in significantly higher database processing activity during inventory-related operations. As inventory data was processed, the increased workload contributed to processing delays, timeout conditions, and degraded application performance for some customers.\n\nTechnical teams developed and deployed hotfixes to address the issue in the affected production environments. Following deployment, teams performed validation activities and monitored impacted processing workflows to confirm recovery.\n\nBy July 18, 2026, at approximately 7:10 AM PDT, validation confirmed successful processing of previously impacted workloads and that the failure patterns associated with the issue were no longer occurring. The incident was considered resolved and technical teams continued monitoring to confirm service stability.  \n\n**Root Cause**\n\nThe incident was caused by a defect introduced in the Flexera One IT Asset Management 2026 R1.1 release affecting inventory-related processing operations.\n\nUnder certain conditions, the defect caused significantly more database processing activity than intended during inventory-related processing operations. This increased workload resulted in elevated resource consumption, processing delays, and timeout conditions.\n\nThese conditions contributed to intermittent application slowness, delayed processing, reconcile failures, and degraded performance affecting some IT Asset Management functionality.  \n\n**Remediation Actions**\n\n1. Incident Investigation Initiated: Technical teams began investigating reports of application slowness, processing delays, and reliability issues affecting APAC and Europe production environments.\n2. Impact Assessment Performed: Teams reviewed customer-reported symptoms, system performance data, and processing behavior to determine the scope and nature of the issue.\n3. Cause Identified: Investigation determined that a defect introduced in the 2026 R1.1 release was generating excessive processing activity during inventory-related operations.\n4. Hotfixes Developed and Deployed: Technical teams developed and deployed hotfixes to correct the affected processing behavior in the APAC and Europe production environments.\n5. Recovery Validated: Teams validated the effectiveness of the hotfixes through monitoring and successful execution of previously impacted processing activities.\n6. Post-Restoration Monitoring Performed: Additional monitoring confirmed that the previously observed failures and performance degradation were no longer occurring and that service stability had been restored.  \n\n**Future Preventative Measures**\n\nBased on the investigation, the following follow-up activity has been identified:\n\n\u2022\tProcessing Logic Improvements: Technical teams implemented corrections to the affected processing logic to eliminate the inefficient behavior that contributed to increased resource consumption and performance degradation.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-18T12:29:05.317-07:00",
"resolved_inferred": false,
"started_at": "2026-07-17T01:24:32.000-07:00",
"state": "postmortem",
"title": "Flexera One- IT Asset management- APAC & EU - Slowness and degraded performance",
"updated_at": "2026-07-30T10:59:28.235-07:00",
"url": "https://stspg.io/s2btch7sfpwk"
},
{
"body": "**Description:** Flexera One - IT Asset Management - EU - Congnos analytics access issue\n\n**Timeframe:**  July 13, 2026, 12:51 PM PDT to July 14, 2026, 1:52 AM PDT\n\n\u200c\n\n**Incident Summary**\n\n\u200c\n\nOn Monday, July 13, 2026, at 12:51 PM PDT, Following the Cognos Analytics 12.1.2 upgrade implemented as part of the Flexera One IT Asset Management release in the Europe \\(EU\\) Production environment, our teams identified an issue that prevented customers from accessing Cognos Analytics. As a result, affected users were unable to successfully launch the application after the scheduled maintenance activities had been completed.\n\nTechnical teams immediately initiated an investigation and confirmed that the issue was isolated to the EU Production environment, with no impact observed in any other production regions. Further analysis determined that the upgrade itself had completed successfully; however, a portion of the required post-upgrade configuration activities had not been fully applied, preventing Cognos Analytics from starting correctly.\n\nThe remaining configuration tasks were promptly completed, restoring normal Cognos Analytics functionality. Following comprehensive validation of customer access and a period of enhanced monitoring, the service was confirmed to be operating normally and the incident was declared resolved.\n\n\u200c\n\n**Root Cause**\n\n\u200c\n\n* Incomplete Post-Upgrade Configuration: Although the Cognos Analytics 12.1.2 upgrade completed successfully, required post-upgrade configuration tasks were not fully completed before the environment was returned to service. \n* Service Initialization Failure: The incomplete configuration prevented Cognos Analytics from initializing correctly, resulting in customers being unable to access the application. \n* Regional Impact: The issue was limited to the EU Production environment. No impact was identified in North America, APAC, or other production regions.\n\n\u200c\n\n**Remediation Actions**\n\n\u200c\n\nTo restore service, technical teams:\n\n* Completed the required post-upgrade configuration activities that had not been fully applied during the deployment.\n* Restarted and validated Cognos Analytics services to confirm successful application initialisation.\n* Performed functional testing and customer access validation to verify that reporting capabilities were operating as expected.\n* Placed the environment under enhanced monitoring following recovery to ensure continued service stability.\n\n\u200c\n\n**Future Preventative Measures**\n\n\u200c\n\n* Enhanced Post-Upgrade Validation: Strengthen deployment validation procedures to verify that all required post-upgrade configuration activities have completed successfully before concluding maintenance activities. \n* Improved Monitoring and Alerting: Enhance Cognos Analytics monitoring and alerting to enable earlier detection of service degradation following upgrades. \n* Deployment Checklist Improvements: Update deployment runbooks and operational checklists to include additional verification of post-upgrade configuration completion prior to returning the environment to production service. \n* Operational Readiness Review: Review upgrade execution procedures to ensure all mandatory post-deployment tasks are completed and validated before maintenance windows are closed.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-14T02:13:56.286-07:00",
"resolved_inferred": false,
"started_at": "2026-07-14T02:13:56.224-07:00",
"state": "postmortem",
"title": "Flexera One - IT Asset Management - EU - Congnos analytics access issue",
"updated_at": "2026-07-26T00:06:54.647-07:00",
"url": "https://stspg.io/grh8pnrkdnnk"
},
{
"body": "**Description:** Flexera One - IT Asset Management - APAC - Missing Menu Items\n\n**Timeframe:** July 12, 2026, 3:00 PM PDT - July 12, 2026, 7:16 PM PDT\n\n**Incident Summary**\n\nOn Sunday, July 12, 2026, at 3:00 PM PDT, customers in the APAC production environment began reporting that several IT Asset Management \\(ITAM\\) menu items, including functionality such as Reports and All Applications, were no longer visible within the Flexera One user interface. Multiple customers were affected, impacting their ability to navigate and access portions of the ITAM application.\n\nTechnical teams immediately began investigating the issue and identified authentication-related errors affecting the ITAM user interface. During the investigation, teams reviewed infrastructure health, application behavior, and recent platform changes to determine the source of the problem.\n\nThe investigation determined that an application configuration change associated with a newly introduced authentication capability had been incorrectly applied to the production environment. The issue became apparent when application instances were refreshed, resulting in authentication-related failures that prevented certain ITAM menu items from loading correctly for affected users.\n\nTechnical teams reverted the affected configuration, refreshed the impacted application instances, and validated recovery across the environment. Following these actions, menu functionality was restored, affected customers confirmed recovery, and the incident was resolved on July 12, 2026, at 7:16 PM PDT after validation and monitoring confirmed normal operation had returned.\n\n**Root Cause**\n\nInvestigation determined that an application configuration change associated with a new authentication capability was incorrectly applied to the production environment.\n\nWhen application instances were subsequently refreshed, the incorrect configuration resulted in authentication-related errors within the ITAM user interface. These errors prevented certain navigation components from loading correctly, causing affected users to experience missing menu items until the configuration was corrected and the affected instances were refreshed.\n\n**Remediation Actions**\n\nThe following actions were taken during the incident response:\n\n1. Incident Investigation Initiated: Technical teams began investigating after receiving reports from multiple customers regarding missing ITAM menu items.\n2. Application Configuration Reviewed: Recent application and configuration changes were reviewed to identify the source of the menu-loading failures.\n3. Incorrect Configuration Identified: Technical teams determined that an authentication-related configuration change had been incorrectly applied to the production environment.\n4. Configuration Corrected: The affected configuration was removed from the impacted production instances.\n5. Application Instances Refreshed: Impacted application instances were refreshed to ensure the corrected configuration was consistently applied across the environment.\n6. Recovery Validation Performed: Technical teams validated menu visibility and functionality across the affected environment and confirmed recovery with impacted customers.\n\n**Future Preventative Measures**\n\nFollowing the incident, corrective measures were implemented to prevent recurrence of the issue and improve detection of similar conditions in the future.\n\n1. Health Check Monitoring Improvements: The existing health check was updated to detect the application behavior associated with this failure condition. The revised monitoring is designed to identify similar menu-loading and authentication-related application failures more effectively and accelerate detection should a similar issue occur in the future.\n2. Deployment Process Reinforcement: The incident highlighted the importance of ensuring new authentication-related features are applied only to their intended environments. The deployment approach and expected application process for these changes have been reinforced with the technical team to reduce the risk of similar configuration issues in the future.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-12T20:44:28.442-07:00",
"resolved_inferred": false,
"started_at": "2026-07-12T16:51:15.759-07:00",
"state": "postmortem",
"title": "Flexera One - IT Asset Management - APAC - Missing Menu Items",
"updated_at": "2026-08-12T12:21:10.427-07:00",
"url": "https://stspg.io/2ly0d68hfzf7"
},
{
"body": "**Description:** Spot Ocean \u2013 AWS \u2013 Ocean ECS Console Performance Degradation\n\n**Timeframe:**  July 6, 2026, 05:45 AM PDT to July 6, 2026, 07:53 AM PDT\n\n\u200c\n\n**Incident Summary**\n\n\u200c\n\nOn Monday, July 6, 2026, at 05:45 AM PDT , our teams identified a performance degradation affecting Spot Ocean ECS for AWS customers. During the impact window, customers experienced increased latency when accessing the Ocean ECS console, and some AWS-related operations were delayed or did not complete successfully.\n\nTechnical teams immediately initiated an investigation to identify the source of the issue. As an AWS service disruption was occurring concurrently in the AWS us-east-1 region, the investigation initially considered both the external AWS event and a recently completed production deployment as potential contributing factors.\n\nThrough detailed analysis, teams determined that the primary cause of the customer impact was instability introduced by the recent Gateway deployment. The resulting degradation in Gateway performance affected request processing and responsiveness within Spot Ocean ECS. While the concurrent AWS service disruption added complexity to the investigation, it was confirmed not to be the primary driver of the customer-facing impact.\n\nTo restore service, teams rolled back the deployment to the previous stable version. Following the rollback, platform performance returned to expected levels and customer workflows resumed normal operation. An extended period of monitoring confirmed sustained service stability before the event was formally resolved.\n\n\u200c\n\n**Root Cause**\n\n\u200c\n\nThe incident was caused by instability introduced in a Gateway deployment .The deployment resulted in degraded performance within the Gateway service, reducing its ability to efficiently process customer requests. This caused increased response times and intermittent failures for Spot Ocean ECS console operations and AWS-related workflows.\n\nRolling back to the previous stable Gateway version restored normal platform performance and resolved the customer impact.\n\nContributing Factors\n\n* A concurrent AWS service disruption in the us-east-1 region occurred during the same timeframe. While it was not the cause of the incident, it complicated the initial investigation and delayed identification of the underlying Gateway issue.\n* The Gateway rollback required additional time due to the scale of the production deployment, extending the overall recovery process.\n* Operations involving resource creation and updates experienced greater impact than read-only activities because of the degraded Gateway responsiveness.\n\n\u200c\n\n**Remediation Actions**\n\n\u200c\n\n* Technical teams investigated the degradation and identified a recent Gateway deployment as the primary source of the issue.\n* The affected Gateway deployment was rolled back to last know stable version, restoring service stability.\n* Platform performance and customer-facing workflows were validated following the rollback.\n* Technical teams completed an extended monitoring period to confirm stable service operation before closing the incident.\n\n\u200c\n\n**Future Preventative Measures**\n\n\u200c\n\n* Improved Deployment Safeguards - Engineering teams will strengthen deployment controls and validation processes to reduce the likelihood of similar deployment-related issues affecting production environments.\n* Enhanced Platform Monitoring - Additional monitoring and alerting will be implemented to provide earlier detection of abnormal Gateway behavior, including service responsiveness, resource utilization, and application stability.\n* Improved Recovery Procedures - Rollback procedures will be enhanced and regularly validated to reduce recovery time and improve operational efficiency during deployment-related incidents.\n* Proactive Service Health Validation - Synthetic health checks will be expanded to continuously validate critical Ocean ECS user workflows, enabling earlier identification of customer-facing performance degradation.\n* Enhanced Dependency Monitoring - Engineering teams will continue improving visibility into external cloud service events and platform dependencies to enable faster differentiation between internal platform issues and third-party service disruptions during future investigations.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-06T09:57:21.570-07:00",
"resolved_inferred": false,
"started_at": "2026-07-06T07:19:02.000-07:00",
"state": "postmortem",
"title": "Spot Ocean \u2013 AWS \u2013 Ocean ECS Console Performance Degradation",
"updated_at": "2026-07-20T00:02:00.757-07:00",
"url": "https://stspg.io/nwr6y4xy9y3n"
},
{
"body": "**Description:** Flexera One \u2013 APAC \u2013 Intermittent Access Issues\n\n**Timeframe:** July 3, 2026, 11:20 AM PDT \u2013 July 3, 2026, 11:50 AM PDT\n\n**Incident Summary**\n\nOn July 3, 2026, at approximately 11:20 AM PDT, an issue was identified affecting a subset of Flexera One services in the APAC production environment.\n\nDuring the incident window, some customers may have experienced intermittent access issues or failures when accessing certain areas of the Flexera One platform. The impact was intermittent in nature, as redundant service capacity remained available and continued serving traffic during the incident. Other Flexera One regions, including NAM and EU, were not impacted.\n\nTechnical teams began investigating immediately and reviewed the affected services and supporting platform components within the APAC environment. The investigation identified that a recent change made within the APAC environment was contributing to the intermittent service failures being observed.\n\nAs part of the recovery effort, the change was reverted and service behavior was monitored to validate recovery. Following the rollback, service stability was restored and customers were able to access affected functionality normally. By approximately 11:50 AM PDT, services were operating as expected and the incident was considered resolved.\n\n\u200c\n\n**Root Cause**\n\nThe incident was caused by a configuration change within the APAC environment that unintentionally affected communication between a subset of platform services and a backend platform component.\n\nDuring the investigation, technical teams determined that the change resulted in certain platform services being unable to communicate properly with the affected component. This led to intermittent failures affecting a subset of Flexera One services in the APAC region.\n\nWhile the issue was occurring, one or more pods supporting the affected services remained available and continued serving traffic. As a result, customers experienced intermittent access issues rather than a complete service outage.\n\n\u200c\n\n**Remediation Actions**\n\nThe following actions were taken during the incident response:    \n\n\u2022\tIncident Investigation Initiated: Technical teams began investigating intermittent access issues affecting a subset of Flexera One services in the APAC region.    \n\u2022\tPlatform Review Completed: Technical teams reviewed the affected services and supporting platform components within the APAC environment to identify the source of the issue.    \n\u2022\tRecent Change Identified: Investigation identified a recently implemented change that correlated with the observed service degradation.    \n\u2022\tChange Reverted: The identified change was reverted as part of the mitigation effort to restore normal service behavior.    \n\u2022\tService Recovery Validated: Technical teams monitored service behavior following the rollback and confirmed that affected functionality was operating normally.    \n\u2022\tPlatform Stability Monitoring Continued: Additional monitoring was performed after restoration activities to verify continued service stability before the incident was closed.\n\n\u200c\n\n**Future Preventative Measures**\n\nThis incident highlighted the importance of validating platform changes to ensure unintended impacts are identified before they affect service availability.    \n\nBased on the investigation, the following follow-up activities are being pursued:    \n\n\u2022\tConfiguration Validation Improvements: Review and enhance validation processes for infrastructure and platform configuration changes to help identify unintended impacts before changes are implemented in production environments.    \n\u2022\tPost-Deployment Verification Enhancements: Review post-deployment verification procedures to ensure critical service dependencies remain accessible and functioning as expected following configuration changes.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-03T12:07:58.110-07:00",
"resolved_inferred": false,
"started_at": "2026-07-03T12:07:58.037-07:00",
"state": "postmortem",
"title": "Flexera One \u2013 APAC \u2013 Intermittent Access Issues",
"updated_at": "2026-07-16T23:12:41.969-07:00",
"url": "https://stspg.io/bwpc9zln0f68"
},
{
"body": "**Description:** Flexera One- IT Asset management- APAC & EU - NDI Inventory Import Failures\n\n**Timeframe:**  June 18, 2026, 09:00 AM PDT to June 25, 2026, 02:42 AM PDT\n\n\u200c\n\n**Incident Summary**\n\n\u200c\n\nOn Wednesday, June 24, 2026, 11:35 PM PDT, following the deployment of the Flexera One IT Asset Management 2026 R1 release to the Australia and Europe Production environments, our teams identified an issue affecting the processing of NDI inventory files. As a result, a significant number of inventory files failed during processing and were not reflected in customer inventory data, causing inventory information displayed in Flexera One ITAM to become outdated for affected customers.\n\nTechnical teams immediately initiated an investigation and determined that the issue occurred within the inventory processing workflow introduced with the 2026 R1 release. Failed inventory files accumulated following repeated processing failures, preventing successful completion of inventory updates and contributing to increased processing load on the affected services.\n\nTo restore service, technical teams implemented corrective actions by refreshing the affected application services across impacted environments. Following the restoration activities, inventory processing resumed successfully, and new inventory uploads were processed as expected. Continued monitoring confirmed stable platform behavior, and the incident was declared resolved after sustained validation.\n\n\u200c\n\n**Root Cause**\n\n\u200c\n\nThe issue was caused by a defect introduced with the Flexera One IT Asset Management 2026 R1 release that affected the processing of NDI inventory files. Under specific conditions, inventory files could not be processed successfully due to an issue within the authentication and processing workflow, causing the files to be rejected before processing could complete.\n\nAs failed files accumulated, processing capacity was increasingly consumed by repeated failures, preventing successful inventory processing and delaying inventory updates for affected customers.\n\nContributing Factors\n\n* Failed inventory files accumulated within the processing workflow, increasing resource utilization on affected inventory services.\n* The increased processing load reduced the capacity available for new inventory processing requests.\n* The issue affected production environments running the 2026 R1 release; North America Production was not impacted because the release had not yet been deployed.\n\n\u200c\n\n**Remediation Actions**\n\n\u200c\n\n* Technical teams identified the affected inventory processing workflow and implemented corrective actions to restore service.\n* Application services were refreshed across the impacted production environments, restoring normal inventory processing.\n* Processing of new inventory files resumed successfully following the service restoration.\n* The environment remained under enhanced monitoring to validate processing performance and confirm sustained recovery.\n\n\u200c\n\n**Future Preventative Measures**\n\n\u200c\n\n* Improved Release Validation - Release validation procedures will be enhanced to include additional end-to-end testing of inventory processing workflows following major platform releases.\n* Enhanced Processing Monitoring - Monitoring and alerting will be reviewed and expanded to provide earlier detection of abnormal inventory processing failures and excessive file accumulation within processing queues.\n* Operational Resiliency - Additional safeguards will be implemented within the inventory processing workflow to improve recovery from processing failures and minimize customer impact should similar conditions occur in the future.\n* Code Quality and Review Improvements \\(Completed\\) - As part of the post-incident review, the affected code has been corrected to ensure failures in communication are handled appropriately and do not result in broader processing impacts. In addition, the development and review process has been reinforced to ensure that review findings are appropriately evaluated and addressed before future changes are approved for release.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-06-24T23:54:13.000-07:00",
"resolved_inferred": false,
"started_at": "2026-06-24T23:54:13.000-07:00",
"state": "postmortem",
"title": "Flexera One- IT Asset management- APAC & EU - Inventory Import Failures",
"updated_at": "2026-07-14T23:19:06.973-07:00",
"url": "https://stspg.io/p5tp7jy2k5q5"
},
{
"body": "**Description:** Flexera One \u2013 EU \u2013 Login Access Disruption\n\n**Timeframe:**  June 24, 2026, 12:41 AM PDT to June 24, 2026, 02:04 AM PDT\n\n\u200c\n\n**Incident Summary**\n\n\u200c\n\nOn Wednesday, June 24, 2026, at 12:41 AM PDT, our teams detected an issue affecting customers in the EU region. Impacted users experienced authentication failures when attempting to sign in to Flexera One using Single Sign-On \\(SSO\\), resulting in server errors and preventing access to the platform.\n\nUpon detection, technical teams immediately initiated an investigation and confirmed that the issue was isolated to the authentication workflow. Validation of the application and underlying platform services showed that all core infrastructure remained healthy and operational. Further analysis determined that a service account used to communicate with the identity provider had become unavailable, causing authentication requests to fail.\n\nThe service account was promptly restored, which re-established normal authentication processing and fully restored customer access. Following restoration, the environment was closely monitored and login functionality was comprehensively validated to ensure service stability. No additional customer impact was observed, and the incident was subsequently declared resolved.\n\n\u200c\n\n**Root Cause**\n\n\u200c\n\nThe issue was caused by the unintended suspension of a service account that supports authentication between Flexera One and the identity provider. As a result, authentication requests could not be processed successfully, preventing affected customers in the EU production region from accessing the platform through Single Sign-On \\(SSO\\).\n\nContributing Factors\n\n* A service account critical to the authentication workflow was inadvertently suspended during a routine cleanup activity.\n* The suspension prevented authentication requests from being validated successfully, resulting in login failures for impacted customers.\n* The issue was isolated to the authentication service and did not affect the availability of the underlying application or platform infrastructure.\n\n\u200c\n\n**Remediation Actions**\n\n\u200c\n\n* The affected service account was restored, re-establishing authentication with the identity provider.\n* Login functionality was validated following restoration.\n* Engineering teams continued monitoring to confirm platform stability and successful customer authentication.\n\n\u200c\n\n**Future Preventative Measures**\n\n\u200c\n\n* Service Account Governance - Processes governing critical service accounts will be strengthened to reduce the risk of inadvertent changes affecting production services.\n* Enhanced Monitoring - Monitoring and alerting will be enhanced to provide earlier detection of authentication failures affecting customer login workflows.\n* Operational Process Improvements - Our teams have documented this as lesson learnt and will implement process improvements to reduce the likelihood of similar incidents.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-06-24T02:00:24.728-07:00",
"resolved_inferred": false,
"started_at": "2026-06-24T01:00:25.516-07:00",
"state": "postmortem",
"title": "Flexera One \u2013 EU \u2013 Login Access Disruption",
"updated_at": "2026-07-07T01:29:41.744-07:00",
"url": "https://stspg.io/xtp281cng8c0"
},
{
"body": "**Description:** Snow Atlas - West Europe - Service Disruption\n\n**Timeframe:**  June 18, 2026, 07:00 AM PDT to June 18, 2026, 08:23 AM PDT\n\n\u200c\n\n**Incident Summary**\n\n\u200c\n\nOn Thursday, 18 June 2026 at 07:00 AM PDT, customers in the West Europe production region experienced a service disruption that affected access to Snow Atlas. During this service event, users encountered HTTP 504 timeout and HTTP 404 errors, which prevented access to the platform and impacted the use of Snow Atlas functionality.\n\nUpon detection, technical teams immediately initiated an investigation and identified the issue within a core platform component responsible for communication between backend services. The degradation disrupted service interactions, resulting in request failures and temporary service unavailability.\n\nThe affected messaging components were restored, enabling dependent services to recover and normal platform operations to resume. Following recovery, extensive validation activities confirmed that service functionality had been fully restored. The environment remained stable under enhanced monitoring, no further customer impact was observed, and the service disruption was formally resolved.\n\n\u200c\n\n**Root Cause**\n\n\u200c\n\nThe service disruption originated from an unexpected failure within a core platform component that facilitates communication between backend services. During the event, multiple messaging components became unavailable simultaneously, preventing critical communication between backend services responsible for processing customer requests.\n\nThe degradation impaired the ability of dependent services to exchange and process requests, leading to timeouts and routing failures. As a result, customers in the affected region experienced difficulty accessing the Snow Atlas platform until service communications were restored and normal operations resumed.\n\nContributing Factors\n\n* Multiple messaging service components became unavailable at the same time, reducing the platform's ability to process inter-service communication.\n* Service communication failures propagated across dependent platform components, resulting in HTTP 504 timeout and HTTP 404 errors.\n* The issue affected the shared messaging infrastructure supporting the West Europe production environment, resulting in widespread customer impact within the region.\n\n\u200c\n\n**Remediation Actions**\n\n\u200c\n\n* Technical teams identified the affected messaging infrastructure components and restored normal service operation.\n* Dependent platform services recovered automatically as messaging functionality was re-established.\n* Service functionality was validated following recovery to confirm successful customer access.\n* Enhanced monitoring was maintained after restoration to verify continued platform stability.\n\n\u200c\n\n**Future Preventative Measures**\n\n\u200c\n\n* Problem Management Review \u2013 A comprehensive review is underway to further validate the root cause and identify long-term corrective actions to prevent recurrence. \n* Platform Resiliency Enhancements \u2013 Opportunities to strengthen the resiliency of the platform's messaging infrastructure will be evaluated to reduce the impact of component-level failures on service availability. \n* Monitoring and Detection Improvements \u2013 Monitoring and alerting capabilities for critical platform components will be enhanced to enable earlier identification of degradation and accelerate response and recovery efforts.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "critical",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-06-18T11:08:29.974-07:00",
"resolved_inferred": false,
"started_at": "2026-06-18T08:33:23.732-07:00",
"state": "postmortem",
"title": "Snow Atlas - West Europe - Service Disruption",
"updated_at": "2026-07-07T01:58:34.713-07:00",
"url": "https://stspg.io/5c0kg3l5cx0m"
},
{
"body": "**Description:** Flexera One \u2013 IT Visibility \u2013 APAC \u2013 Data Explorer Service Disruption\n\n**Timeframe:** June 17, 2026, 6:47 PM PDT \u2013 June 18, 2026, 7:46 AM PDT\n\n**Incident Summary**\n\nDuring the incident window, customers experienced issues with Data Explorer functionality for Flexera One IT Visibility in the APAC region.\n\nAffected customers experienced failures when running queries, delays and errors when retrieving data, and issues accessing related functionality that depends on Data Explorer. Some dependent functionality, including reporting and connections, did not operate as expected.\n\nTechnical teams investigated and identified that Data Explorer and related services were unable to successfully communicate with required internal services used for request processing. This resulted in query execution failures and degraded behavior for related data and connection workflows.\n\nAs part of the recovery effort, technical teams identified the issue and reverted the changes to restore the previous stable state. Following the rollback, connectivity between services was restored, and Data Explorer, query execution, reporting, connections, and related functionality returned to expected behavior.\n\nBy June 18, 2026, at approximately 7:46 AM PDT, services had been validated and confirmed to be operating normally across APAC. The impact was limited to APAC. North America and Europe remained unaffected throughout the incident.\n\n**Root Cause**\n\nThe issue occurred during an ongoing migration of internal platform services, where services were being moved incrementally to a new platform. During this activity, a change introduced while migrating a majority of dependent services caused a disruption in service-to-service connectivity. This impacted downstream components that rely on those services for request processing.\n\nThe issue was not observed during validation performed before and after the migration activity. After the connectivity disruption occurred, Data Explorer and related services were unable to successfully communicate with the required internal services. This caused failures when running queries, delays and errors when retrieving data, and issues with dependent functionality including reporting and connections.\n\n**Remediation Actions**\n\nThe following actions were taken during the incident response:\n\n1. **Incident Investigation Initiated:** Technical teams investigated reports of Data Explorer issues affecting Flexera One IT Visibility in the APAC region.\n2. **Impact Scope Reviewed:** The issue was confirmed to affect APAC. North America and Europe were reviewed and confirmed to be operating as expected.\n3. **Dependent Functionality Reviewed:** Technical teams reviewed related functionality, including query execution, reporting, and connection-related workflows.\n4. **Service Communication Issue Identified:** Technical teams identified that downstream services were unable to successfully communicate with required internal services.\n5. **Migration Activity Reviewed:** Technical teams determined that the issue occurred during an ongoing migration of internal platform services to a new platform.\n6. **Changes Reverted:** The changes introduced during the migration activity were reverted to restore the previous stable state.\n7. **Service Connectivity Restored:** Following the rollback, connectivity between services was restored.\n8. **Service Recovery Validated:** Technical teams validated Data Explorer, query execution, reporting, connections, and related functionality and confirmed that services were operating normally in APAC.\n\n**Future Preventative Measures**\n\nThis incident highlighted the importance of planning, visibility, and validation for internal platform migration activities that may affect customer-facing services or downstream dependencies.\n\nBased on the investigation, the following follow-up activities are being reviewed:\n\n* **Migration Assessment and Planning:** Review assessment and planning practices for internal platform migration activities to better identify downstream dependencies, required validation steps, impacted workflows, and any potential customer impact before migration work is performed.\n* **Scheduled Maintenance Planning:** Ensure internal platform migration activities with any potential customer impact are planned and communicated through a scheduled maintenance window.\n* **Validation Coverage Review:** Review validation practices for migration activity to help ensure testing covers dependent customer-facing workflows before and after migration work is performed, especially where prior validation does not identify service-to-service connectivity issues.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-06-18T08:06:14.296-07:00",
"resolved_inferred": false,
"started_at": "2026-06-18T07:44:23.000-07:00",
"state": "postmortem",
"title": "Flexera One - IT Visibility - APAC - Data Explorer Service Disruption",
"updated_at": "2026-07-06T23:43:27.604-07:00",
"url": "https://stspg.io/5wbphy2y0qzn"
},
{
"body": "**Description:** Flexera One- IT Asset management- APAC - Page loading issues\n\n**Timeframe:**  June 18, 2026, 05:22 AM PDT to June 18, 2026, 06:09 AM PDT\n\n\u200c\n\n**Incident Summary**\n\n\u200c\n\nOn Thursday, 18 June 2026, at 05:22 AM PDT, our teams detected an issue affecting customers in the APAC production region. Impacted users were unable to access the Flexera One IT Asset Management \\(ITAM\\) service, as the ITAM landing page failed to load and key navigation components did not initialise correctly. As a result, customers were unable to access ITAM functionality through the user interface.\n\nUpon detection, our technical teams immediately launched an investigation and confirmed that the issue was isolated to the APAC region. Initial assessments indicated that the underlying application infrastructure remained healthy, which focused the investigation on the authentication and service authorisation components supporting the platform.\n\nFurther analysis determined that a recent configuration change introduced during a service migration resulted in authentication failures between internal platform services. Although pre-migration testing indicated no expected service impact, the configuration behaved unexpectedly under specific conditions, preventing successful service-to-service authentication. As a result, a critical backend service responsible for tenant management became unavailable. This prevented the ITAM application from retrieving the tenant information required to initialise the user interface, causing customers to be unable to access the platform.\n\nOnce the configuration issue was identified, the change was rolled back to restore normal service communication. Following the rollback, technical teams validated the functionality of all affected services and confirmed that customer access had been fully restored. The environment remained stable throughout the post-restoration monitoring period, and the incident was formally declared resolved.\n\n\u200c\n\n**Root Cause**\n\n\u200c\n\nThe incident was triggered by a configuration change introduced during a planned service migration. Based on pre-migration testing, no service disruption was anticipated. However, under specific conditions, the configuration prevented successful authentication between internal platform services, resulting in the unavailability of a critical backend service responsible for tenant management. As a consequence, the ITAM application was unable to retrieve the tenant information required to initialise the user interface, preventing customers from accessing the platform.\n\nContributing Factors\n\n* The configuration change affected authentication between internal platform services in the APAC environment.\n* A backend service supporting tenant management was unable to authenticate with its underlying data store because of the configuration change.\n* Since the ITAM user interface depends on successful initialization of this service, customers experienced missing navigation elements and page load failures.\n* The issue was isolated to the APAC region, as the configuration change was specific to that environment.\n\n\u200c\n\n**Remediation Actions**\n\n\u200c\n\n* Technical teams identified the configuration change responsible for the authentication failures.\n* The configuration was rolled back, restoring normal communication between platform services.\n* Service functionality and customer access were validated following the rollback.\n* Enhanced monitoring was maintained to confirm continued platform stability after restoration.\n\n\u200c\n\n**Future Preventative Measures**\n\n\u200c\n\n* Strengthened Configuration Validation - Configuration changes associated with infrastructure migrations will undergo additional validation and production readiness checks before deployment.\n* Enhanced Monitoring and Alerting - Monitoring capabilities will be expanded to provide earlier detection of authentication and service authorization failures affecting critical platform components.\n* Improved Change Governance - The change management process for infrastructure migrations will be enhanced with additional validation checkpoints and post-deployment verification to reduce the likelihood of similar configuration-related issues in the future.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "critical",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-06-18T06:16:11.439-07:00",
"resolved_inferred": false,
"started_at": "2026-06-18T05:45:37.467-07:00",
"state": "postmortem",
"title": "Flexera One- IT Asset management- APAC - Page loading issues",
"updated_at": "2026-07-01T22:51:51.801-07:00",
"url": "https://stspg.io/xcwwv2x5mzgw"
},
{
"body": "**Description:** Snow Atlas - Australia - Few SAM core pages are down\n\n**Timeframe:**  June 11, 2026, 05:44 PM PDT to June 11, 2026, 11:22 PM PDT\n\n\u200c\n\n**Incident Summary**\n\n\u200c\n\nOn June 11, 2026, a small subset of customers in the Australia production environment began experiencing issues accessing Snow Atlas SAM Core functionality. Affected users encountered HTTP 500 errors, application page failures, or pages that remained in a continuous loading state. The impact was limited to SAM Core functionality, while SaaS Manager remained fully operational.\n\nTechnical teams immediately initiated an investigation and determined that the issue was isolated to the affected tenant environments. The investigation identified that a critical database service required to support SAM Core operations was not running following routine infrastructure maintenance. As a result, several backend processes required by the application were unavailable, leading to the observed customer impact.\n\nOur teams restored the affected database service across all impacted environments, after which normal application functionality resumed. Validation confirmed that service had been restored for all affected customers, and the environment remained stable throughout the monitoring period.\n\n\u200c\n\n**Root Cause**\n\n\u200c\n\nThe issue was caused by a configuration that had been introduced during a previous migration between Australia regions. The configuration was intended to support the migration process temporarily but was not reverted once the migration had been completed.\n\nAlthough the configuration had no immediate operational impact, it affected the behavior of a critical database service during a routine infrastructure maintenance activity on June 11. When infrastructure components were restarted as part of planned maintenance, the database service did not automatically recover in the affected environments. This prevented several backend processes supporting SAM Core functionality from operating correctly, resulting in application errors for affected customers.\n\nContributing Factors\n\n* A temporary migration configuration remained in place after the migration had been completed.\n* Routine infrastructure maintenance exposed the latent configuration issue.\n* A critical database service did not automatically recover following the infrastructure restart.\n* Existing monitoring did not immediately detect the service recovery failure, delaying identification of the issue.\n\n\u200c\n\n**Remediation Actions**\n\n\u200c\n\n* Technical teams identified the affected database service and restored normal service operation across all impacted environments.\n* Customer functionality was validated following recovery to confirm that affected SAM Core pages were operating normally.\n* The temporary migration configuration has been corrected to prevent recurrence.\n* The environment was monitored following restoration to confirm continued platform stability.\n\n\u200c\n\n**Future Preventative Measures**\n\n\u200c\n\n* Improved Configuration Governance - Migration-related configuration changes will be subject to enhanced post-migration validation to ensure temporary settings are removed once migration activities have been completed.\n* Enhanced Service Monitoring - Monitoring and alerting will be reviewed to provide earlier detection of failures affecting critical backend services required for customer-facing functionality.\n* Operational Process Improvements - Migration and maintenance procedures will be updated to incorporate the lessons learned from this issue, reducing the likelihood of similar issues occurring in future maintenance activities.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-06-11T23:24:20.947-07:00",
"resolved_inferred": false,
"started_at": "2026-06-11T18:38:18.099-07:00",
"state": "postmortem",
"title": "Snow Atlas - Australia - Few SAM core pages are not loading",
"updated_at": "2026-06-25T22:31:45.937-07:00",
"url": "https://stspg.io/fw1q5wklzkf3"
},
{
"body": "**Description:** Flexera One - IT Asset Management - All Regions - Third-Party Inventory Import Issues\n\n**Timeframe:**  June 4, 2026, 1:45 PM PDT to June 08, 2026, 00:47 PDT \n\n\u200c\n\n**Incident Summary**\n\n\u200c\n\nOn Thursday, June 4, 2026, at 1:45 PM PDT, our teams identified an issue affecting third-party inventory imports in IT Asset Management caused some inventory uploads to time out, fail, or remain in progress longer than expected. The issue impacted the processing of certain third-party inventory sources and resulted in delays to inventory data availability for affected customers.\n\nInvestigation determined that, following a recent platform change, a subset of third-party inventory imports encountered issues during the transfer of inventory data into IT Asset Management. \n\nThe issue was initially identified through a combination of customer reports and internal monitoring. Further investigation confirmed that while inventory data was successfully generated and processed through earlier stages of the workflow, certain uploads failed to complete successfully, resulting in import processing delays and failures. Due to the nature of the response received during these failures, the underlying cause was not immediately apparent, which extended the time required for diagnosis. Technical teams performed a detailed analysis of the processing workflow and identified that certain larger inventory uploads were not being handled as expected under specific conditions. A solution was implemented to improve how these larger data transfers are managed.\n\nThe corrective update was validated in a controlled environment, followed by comprehensive end\u2011to\u2011end testing, and then deployed to production. Following deployment, inventory processing returned to expected behaviour by Jun 06, 2026 , 01:03 PDT. Continued monitoring confirmed stable performance, and no further recurrence of the issue was observed. After extended monitoring, the incident was declared as resolved on Jun 08, 2026 at 00:47 PDT.\n\n\u200c\n\n**Root Cause**\n\n\u200c\n\nThe issue was caused by limitations encountered during the transfer of larger inventory data sets between platform services following the production change. Under specific conditions, larger inventory uploads were not processed successfully, leading to incomplete or delayed processing. When these uploads failed, the receiving service returned a generic client error response rather than an error specifically indicating a file size limitation. As a result, the sending service interpreted the response as a capability or configuration issue rather than an upload failure, preventing the true cause from being immediately identified.\n\nContributing Factors\n\n* The response returned during failure scenarios did not clearly indicate the underlying condition.\n* This response behaviour initially led to misinterpretation of the failure, extending diagnosis time.\n* Pre release validation included large data scenarios; however, production conditions introduced additional variability not fully represented during testing.\n* Platform configuration indicated support for larger data handling, which made the practical limitation difficult to anticipate during design and validation.\n\n\u200c\n\n**Remediation Actions**\n\n\u200c\n\n* The processing workflow was analysed to identify the point of failure within the inventory transfer process. \n* Data transfer behaviour between platform components was reviewed and validated. \n* A revised approach was implemented to transfer larger data sets in smaller segments. The solution was validated through staging and end to end testing to ensure data integrity and processing completion. \n* The validated fix was deployed to production environments. \n* Enhanced monitoring was maintained throughout rollout and validation to confirm successful processing and overall stability.\n\n\u200c\n\n**Future Preventative Measures**\n\n\u200c\n\n* Improved Data Transfer Handling - Large inventory data transfers configuration has been updated, improving reliability and reducing dependency on single transaction limits.\n* Enhanced Error Reporting - Improvements will be made to ensure that system responses more clearly reflect the underlying cause of failures, enabling faster and more accurate diagnosis.\n* Expanded Validation Scenarios - Additional testing scenarios involving larger and more complex data sets will be incorporated to better reflect real world conditions prior to release.\n* Stronger Integration Validation - Future changes will include deeper validation of interactions between platform components to identify and mitigate potential constraints earlier.\n* Monitoring and Detection Enhancements - Monitoring and alerting capabilities will be further enhanced to provide earlier visibility into processing delays or failures, helping to reduce time to detection and resolution.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-06-08T00:47:00.117-07:00",
"resolved_inferred": false,
"started_at": "2026-06-04T14:17:22.000-07:00",
"state": "postmortem",
"title": "Flexera One - IT Asset Management - All Regions - Third-Party Inventory Import Issues",
"updated_at": "2026-06-18T23:39:54.334-07:00",
"url": "https://stspg.io/zfxrlqfn6rdd"
},
{
"body": "**Description:** Flexera One IT Asset Management \\(ITAM\\) - North America - Delayed Batch Processing\n\n**Timeframe:** June 2, 2026, 10:29 PM PDT - June 4, 2026, 6:29 AM PDT  \n\n**Incident Summary**  \nOn June 2, 2026, following a scheduled production release in the North America production environment, customers using Flexera One IT Asset Management \\(ITAM\\) experienced delays in batch processing activities.\n\nAfter the release completed, technical teams identified that batch processing workloads were accumulating in processing queues and were not being executed at the expected rate. As the backlog increased, customers experienced delays in processing activities that relied on the ITAM batch processing platform.\n\nTechnical teams immediately began investigating the issue and observed that processing capacity was not being utilized as expected. Additional batch processing capacity was temporarily introduced, workloads were reviewed and rebalanced, and non-critical processing tasks were reduced to help restore overall processing throughput. These actions resulted in a steady improvement in processing rates and a gradual reduction of the accumulated backlog.\n\nThroughout the recovery effort, technical teams closely monitored queue levels and processing activity while maintaining additional capacity to accelerate backlog reduction. Processing throughput continued to improve and queued workloads steadily decreased until normal operating conditions were restored.\n\nBy June 4, processing backlogs had been cleared, remaining queued work was processing normally, and technical teams confirmed there was no longer any customer-facing impact. Following validation and continued monitoring, the incident was resolved.  \n\n**Root Cause**\n\nInvestigation determined that a release pipeline sequencing issue allowed infrastructure deployment activities to begin before a required release step had been completed.\n\nAs a result, server instances became available earlier than intended during the production release process. This created conditions that led to abnormal batch processing behavior and the accumulation of processing workloads within the batch queues.\n\nThe resulting backlog significantly reduced overall processing throughput and delayed execution of customer batch processing activities until mitigation measures were implemented and normal processing rates were restored.  \n\n**Remediation Actions**\n\nThe following actions were taken during the incident response:\n\n1. Incident Investigation Initiated: Technical teams began investigating after identifying abnormal growth in batch processing queues following the production release.\n2. Additional Processing Capacity Added: Additional batch processing instances were deployed to increase processing throughput and accelerate backlog reduction.\n3. Workload Rebalancing Performed: Processing workloads were reviewed and rebalanced to improve utilization of available processing resources.\n4. Queue Cleanup Activities Executed: Non-critical processing workloads were reduced to allow processing resources to focus on customer-impacting batch activities and improve overall processing throughput.\n5. Continuous Monitoring and Validation: Technical teams continuously monitored queue levels, processing throughput, and backlog reduction until normal operating conditions were restored.\n6. Release Pipeline Corrected: The release pipeline was updated to ensure required release steps are completed before infrastructure deployment activities can proceed.  \n\n**Future Preventative Measures**\n\nFollowing the incident, changes were implemented to the release process to prevent the condition that contributed to this issue from occurring during future deployments.\n\n1. Release Process Improvement: The release process was updated to ensure that critical deployment steps are completed in the correct order before server deployment activities begin. This prevents server components from starting earlier than intended and reduces the risk of similar batch processing disruptions during future releases.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-06-04T11:58:48.893-07:00",
"resolved_inferred": false,
"started_at": "2026-06-03T09:00:59.654-07:00",
"state": "postmortem",
"title": "Flexera One IT Asset Management - NA - Batch Job Processing Delays",
"updated_at": "2026-08-11T19:24:05.436-07:00",
"url": "https://stspg.io/wrm285d55kbq"
},
{
"body": "**Description:** Snow Atlas - West EU - Service Disruption\n\n**Timeframe:** June 3, 2026, 06:33 AM PDT to June 3, 2026, 07:28 AM PDT\n\n\u200c\n\n**Incident Summary**\n\n\u200c\n\nOn Wednesday,  June 3, 2026 at 06:33 AM PDT, customers in the West EU production environment experienced a service disruption that prevented access to Snow Atlas through snowsoftware.io. Affected users were unable to log in to the platform and, in some cases, encountered service availability errors during authentication attempts.\n\nUpon investigation, our teams determined that a critical identity and authentication service within the West EU environment became unavailable during infrastructure maintenance activities being performed in a separate region. As a result, authentication requests could not be processed successfully, preventing customers from accessing the platform.\n\nOnce the issue was identified, the affected service was promptly restored, and normal authentication functionality resumed. Following restoration, technical teams conducted validation activities and continued enhanced monitoring to confirm platform stability and verify that customer access had been fully restored.\n\n\u200c\n\n**Root Cause**  \n\nThe issue was caused by an operational error during planned maintenance activities. While performing maintenance preparations in a separate production region, an action intended for that environment was inadvertently executed against the West EU environment. This resulted in a critical identity and authentication service being unintentionally scaled down, causing authentication failures and preventing customer access to Snow Atlas. \n\n\u200c\n\n**Contributing Factors**\n\n\u200c\n\n* A delay in updating the active management session resulted in commands being executed against the unintended environment.\n* The affected authentication service represented a critical dependency for customer login and platform access.\n* Existing monitoring and alerting did not provide timely notification to the responsible engineering team when the authentication service became unavailable, extending the time required to identify and remediate the issue.\n\n\u200c\n\n**Remediation Actions**\n\n\u200c\n\nThe following remediation steps were implemented to restore service functionality:\n\n* The affected authentication service was restored, re-establishing customer access to the platform.\n* Technical teams validated service functionality and confirmed successful customer authentication following recovery.\n* Engineering teams reviewed maintenance procedures and execution logs to identify the sequence of events that led to the incident.\n* Monitoring was maintained following restoration to verify continued platform stability and service availability. \n\n\u200c\n\n**Future Preventative Measures**\n\n* Enhanced Change Controls - Maintenance procedures will be updated to include additional safeguards and verification steps before executing operational actions across production environments.\n* Strengthened Monitoring and Alerting - Monitoring and alerting coverage for critical authentication services will be enhanced to ensure responsible teams receive immediate notification of service degradation or outages.\n* Operational Process Improvements - The lessons learned from this incident have been incorporated into our maintenance and change management practices. These improvements will further strengthen operational controls, reduce the likelihood of similar events, and improve our ability to detect and respond to service disruptions.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "critical",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-06-03T07:49:56.297-07:00",
"resolved_inferred": false,
"started_at": "2026-06-03T07:49:56.243-07:00",
"state": "postmortem",
"title": "Snow Atlas - West EU - Service Disruption",
"updated_at": "2026-06-16T01:37:29.491-07:00",
"url": "https://stspg.io/mndtgdwh1k4x"
},
{
"body": "**Description:** Snow Atlas - Australia - Service Unavailable\n\n**Timeframe:** June 02, 2026, 5:30 AM PDT to June 02, 2026, 9:18 AM PDT\n\n\u200c\n\n**Incident Summary**\n\n\u200c\n\nOn Tuesday, June 02, 2026, 5:30 AM PDT, A service disruption affected Snow Atlas customers in the Australia region following a scheduled maintenance activity. As part of a planned infrastructure maintenance activity, our teams performed a migration between Australia South East and Australia East. The maintenance was expected to be completed within the scheduled window; however, the activity extended beyond the planned duration due to unexpected delays encountered during execution.\n\nThroughout the issue, customers retained access to the Snow Atlas application and were able to view and browse existing data. The impact was limited to data ingestion and data export functionality. New ingestion requests and export requests were queued while maintenance activities remained in progress.\n\nDuring the maintenance execution, technical teams encountered delays related to deployment synchronization activities and infrastructure management tooling, which extended the overall duration of the migration. Once the migration was completed and services were re-enabled, queued ingestion and export requests resumed processing normally by 9:18 AM PDT.\n\n\u200c\n\n**Root Cause**\n\n\u200c\n\nThe issue was caused by a scheduled infrastructure migration that exceeded its planned maintenance window. The migration required multiple coordinated deployment, configuration, and application management activities, each dependent on successful completion of the previous step.\n\nDuring execution, several deployment and synchronization operations took significantly longer than anticipated based on testing and pre-maintenance validation. These delays extended the maintenance activity beyond the expected timeframe, resulting in a longer-than-planned interruption to data ingestion and export services.\n\nContributing Factors\n\n* The migration involved multiple interdependent deployment and configuration changes that required sequential execution.\n* Deployment synchronization activities experienced intermittent delays and required additional time to complete.\n* Certain deployment operations required manual intervention when synchronization did not complete as expected.\n* The duration of these activities was longer than observed during pre-production testing, resulting in an inaccurate estimate of the total maintenance execution time.\n\n\u200c\n\n**Remediation Actions**\n\n\u200c\n\nThe following remediation steps were implemented to restore service functionality:\n\n* Technical teams completed the migration between Australia South East and Australia East.\n* Deployment and synchronization issues encountered during the migration were investigated and resolved.\n* Data ingestion services were re-enabled following successful completion of the migration activities.\n* Queued ingestion and export requests were allowed to process normally once services were restored.\n* The environment was closely monitored following restoration to verify platform stability and successful processing of queued requests.\n\n\u200c\n\n**Future Preventative Measures**\n\n* Improved Maintenance Planning - Future migration activities will be further segmented where possible to reduce the number of changes performed within a single maintenance window and improve execution predictability.\n* Enhanced Maintenance Duration Estimation - Maintenance planning processes will be updated to better account for deployment synchronization times, infrastructure dependencies, and potential manual intervention requirements when estimating maintenance windows.\n* Lessons Learned Review - The maintenance execution process and associated dependencies will be reviewed to identify additional opportunities to streamline migration activities, reduce operational complexity, and minimize the risk of maintenance overruns in future events.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-06-02T09:18:58.000-07:00",
"resolved_inferred": false,
"started_at": "2026-06-02T07:46:38.000-07:00",
"state": "postmortem",
"title": "Snow Atlas - Australia - Data Ingestion and Export Disruption",
"updated_at": "2026-06-16T23:25:22.257-07:00",
"url": "https://stspg.io/q4jb9f5rg6p2"
},
{
"body": "**Description:** Flexera One - IT Visibility - NA - Inventory Processing Delays and Timeout Status\n\n**Timeframe:** June 2, 2026, 7:29 AM PDT - June 15, 2026, 7:58 AM PDT\n\n**Incident Summary**\n\nOn June 2, 2026, at approximately 7:29 AM PDT, we received customer reports of IT Visibility inventory imports displaying timeout status in the North America production environment.\n\nThe ITV platform remained accessible during the incident. The primary customer impact was delayed inventory processing and delayed availability of updated inventory data for some customers. In some cases, imports displayed timeout status while processing was still pending or delayed. This status did not always indicate that processing had failed, as some processing activity could continue in the background and complete later.\n\nTechnical teams investigated reported customer examples, reviewed organization and data source-level processing activity, and monitored backlog and processing progress over time. During the investigation, teams identified multiple contributing factors affecting timely inventory processing and recovery. These included processing inefficiencies that could add unnecessary load, conditions where processing progress could be delayed or regress during recovery, and temporary database instability that affected backlog reduction.\n\nCorrective actions were applied to improve processing behavior, stabilize recovery, and unblock affected processing activity. Technical teams continued monitoring timed-out imports, customer-reported examples, and new processing activity to confirm recovery.\n\nBy June 15, 2026, at approximately 7:58 AM PDT, the broader incident impact had cleared. Remaining isolated items were isolated from the broader incident impact and continued to be tracked separately through follow-up actions.\n\n**Root Cause**\n\nThe incident was caused by multiple contributing factors affecting the timely processing and availability of updated IT Visibility inventory data.\n\nSome inventory imports experienced delayed completion due to processing backlog and inefficiencies within the inventory processing flow. Identified defects contributed to unnecessary processing load and, in some cases, caused recovery progress to regress when services were restarted.\n\nTechnical teams also identified conditions where queued processing activity was not always progressing as expected without recovery actions. In addition, temporary database instability impacted overall service recovery and slowed backlog reduction while the system was catching up.\n\nAs a result, some imports displayed timeout status when processing took longer than expected. In some cases, the timeout status did not mean processing had stopped or failed, as processing could continue in the background and complete later.\n\n**Remediation Actions**\n\nThe following actions were taken during the incident response:\n\n1. **Incident Response Initiated:** Technical teams began investigating after multiple customer reports were received regarding inventory imports displaying timeout status.\n2. **Customer Reports Reviewed:** Reported customer examples and related support cases were reviewed to validate the scope, affected organizations, data sources, and current processing behavior.\n3. **Regional Scope Validated:** Technical teams reviewed timeout activity across regions. APAC did not show timeout activity, EU impact was limited and later narrowed, and the primary ongoing impact remained in North America.\n4. **Processing Activity and Backlog Reviewed:** Teams reviewed inventory processing activity, backlog queues, organization-level processing state, and data source-level status to determine where imports were delayed and whether processing was still progressing.\n5. **Database Recovery Supported:** Technical teams worked through database instability affecting recovery, including cluster recovery and scaling activity, to restore stability and improve backlog reduction.\n6. **Processing Fixes Applied:** Corrective fixes were deployed for identified processing defects that could contribute to timeout behavior, unnecessary processing load, and recovery progress moving backwards during service restarts.\n7. **Queued Processing Recovery Performed:** Teams applied recovery actions and workarounds to help queued processing activity move forward while the remaining processing issue was investigated and addressed.\n8. **Affected Imports Unblocked:** Technical teams implemented a workaround to unblock affected imports for the remaining impacted organizations and continued monitoring those imports through completion.\n9. **Customer Examples Validated:** Customer-reported examples were reviewed throughout recovery to confirm whether affected imports and updated inventory data were progressing as expected.\n10. **Post-Recovery Monitoring Performed:** Teams continued monitoring platform activity, timed-out imports, and new processing behavior until the broader incident impact had cleared. Remaining isolated items were isolated from the broader incident impact and continued to be tracked separately through follow-up actions.\n\n**Future Preventative Measures**\n\nFollowing this incident, technical teams implemented improvements to strengthen the reliability, recoverability, monitoring, and operational visibility of IT Visibility inventory processing.\n\nThe implemented improvements included:\n\n1. **Inventory Processing Reliability and Recoverability:** Improvements were made to reduce conditions that could contribute to delayed processing, unnecessary processing load, or recovery delays when queued inventory activity needed to catch up.\n2. **Monitoring and Observability:** Additional monitoring and observability improvements were implemented to provide better visibility into processing health, throughput, latency, failure patterns, retry behavior, and backlog accumulation across production environments.\n3. **Recovery Handling:** Recovery processes were improved to help teams identify delayed or stuck processing activity earlier and take corrective action more consistently when processing did not progress as expected.\n4. **Operational Readiness:** Internal investigation and validation processes were improved to help teams assess affected organizations, data sources, and processing state more efficiently during similar events.\n\nTeams also reviewed import status behavior to provide clearer status information when processing takes longer than expected but may still continue in the background.\n\nSince these corrective measures were implemented, we have not observed a recurrence of the same broad incident pattern.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-06-16T02:33:49.129-07:00",
"resolved_inferred": false,
"started_at": "2026-06-02T07:29:05.000-07:00",
"state": "postmortem",
"title": "Flexera One - IT Visibility - NA & EU - Inventory Processing Delays and Timeout Status",
"updated_at": "2026-08-10T18:11:42.246-07:00",
"url": "https://stspg.io/5dt5dl6jg11s"
},
{
"body": "**Description:** Flexera One \u2013 IT Asset Management and IT Visibility \u2013 APAC \u2013 Beacon Configuration and Reconfiguration Errors\n\n**Timeframe:** June 1, 2026, 7:41 PM PDT \u2013 June 2, 2026, 12:58 PM PDT  \n\n**Incident Summary**\n\nOn Monday, June 1, 2026, at approximately 7:41 PM PDT, an issue was identified affecting Beacon configuration and reconfiguration workflows in the APAC Production environment for Flexera One IT Asset Management and IT Visibility.\n\nDuring the affected period, customers attempting to access affected Beacon-related pages may have encountered a red bar error or server-side application error.\n\nTechnical teams investigated the impacted workflows and backend service dependencies. The investigation confirmed that affected requests were returning server-side errors and that the issue was associated with the backend service supporting Beacon-related workflows.\n\nThe issue was resolved by updating the affected Beacon API service in APAC Production. Following the update, the affected Beacon pages were confirmed accessible again. By Tuesday, June 2, 2026, at approximately 12:58 PM PDT, service was confirmed restored and the incident was considered resolved.  \n\n**Root Cause**\n\nThe incident was caused by a compatibility issue between the deployed Beacon API service version in APAC Production and recent IT Asset Management database changes. A required related Beacon API service update was not completed in APAC Production as part of the related release activity, which caused some Beacon configuration and reconfiguration workflows to return backend errors.\n\nThese backend errors were surfaced in the UI as red bar errors when customers attempted to access affected Beacon-related pages.  \n\n**Remediation Actions**\n\nThe following actions were taken during the incident response:\n\n* **Impact Assessment:** Technical teams confirmed the issue was affecting Beacon configuration and reconfiguration workflows in APAC Production.\n* **Backend Error Isolation:** Investigation confirmed that affected requests were returning server-side errors from backend service dependencies.\n* **Service Compatibility Review:** Technical teams identified a compatibility issue between the deployed Beacon API service version and recent IT Asset Management database changes.\n* **Corrective Service Update:** The Beacon API service was updated in APAC Production to restore the affected workflows.\n* **Service Restoration Validation:** Following the update, technical teams validated that the affected Beacon pages were accessible and functioning as expected.  \n\n**Future Preventative Measures**\n\nThe following follow-up actions were identified:\n\n* **Strengthened Deployment Procedures:** We are strengthening deployment dependency checks and cross-team release coordination to ensure required service updates are clearly identified, tracked, and completed when related database or platform changes are released.\n* **Regional Version Alignment:** Related Beacon API service updates were completed across additional regions, including EU and NAM, to maintain version consistency and reduce the risk of similar compatibility issues occurring elsewhere.\n* **Expanded Post-Release Validation:** We are expanding post-release validation for critical Beacon configuration and reconfiguration workflows, especially when backend service or database changes are involved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-06-02T14:55:49.583-07:00",
"resolved_inferred": false,
"started_at": "2026-06-02T00:41:12.788-07:00",
"state": "postmortem",
"title": "Flexera One- IT Asset management & IT Visibility- APAC- Red Bar Error When Reconfiguring A Beacon",
"updated_at": "2026-06-16T22:22:05.690-07:00",
"url": "https://stspg.io/kqrnkyqml65h"
},
{
"body": "**Description:** Snow Atlas \u2013 APAC \u2013 Login Failures and HTTP 500 Errors\n\n**Timeframe:** May 28, 2026, 20:57 PDT to May 29, 2026, 03:32 PDT\n\n\u200c\n\n**Incident Summary**\n\nOn Thursday, May 28, 2026, at 20:57 PDT, an incident was identified impacting a subset of Snow Atlas customers in the APAC production environment. Affected users were unable to access the platform or experienced degraded application functionality. This included HTTP 500 and 404 errors, as well as instances where pages loaded incompletely with no data before ultimately returning service unavailability messages.\n\nTechnical teams engaged immediately and identified degraded performance within a core routing component responsible for handling traffic. As the issue progressed, backend services were unable to communicate reliably, leading to increased request failures and overall application instability for impacted tenants.\n\nInitial mitigation actions, including targeted service restarts, successfully restored access for affected customers. Teams continued to investigate the underlying cause while maintaining enhanced monitoring to validate platform stability. A subsequent controlled restart and recovery of the impacted routing services fully resolved the contention conditions.\n\nFollowing these actions, no further customer impact was observed. The platform remained stable throughout the monitoring period, and the incident was formally resolved after normal service operations were confirmed.\n\n\u200c\n\n**Root Cause**\n\nThe incident was triggered by an automated infrastructure upgrade initiated by the cloud service provider. This upgrade resulted in a sequential restart of infrastructure nodes supporting the Snow Atlas platform, causing a large number of application services to restart within a compressed timeframe.\n\nDuring the recovery phase, a critical routing component responsible for service discovery and request routing became saturated due to the high volume of service registration activity. Consequently, routing information was intermittently unavailable to dependent services, leading to application errors, request failures, and overall service instability for a subset of APAC customers.\n\nContributing Factors:\n\n* The provider-initiated infrastructure upgrade was unplanned from an application perspective and triggered a large-scale, simultaneous service restart across the environment.\n\n* The routing service had limited redundancy, increasing the potential for service disruption under certain conditions.\n\n* During recovery, the temporary unavailability of the routing service prevented dependent services from resolving required routes, resulting in HTTP 404 and HTTP 500 errors across multiple application functions.\n\n* Elevated recovery activity, combined with the lack of redundancy, amplified both the impact and duration of the incident.\n\n\u200c\n\n**Remediation Actions**\n\nThe following remediation steps were implemented to restore service functionality:\n\n* Technical teams identified the affected routing component and performed targeted recovery actions to restore service.\n\n* Service restarts were completed to re-establish routing functionality and stabilize affected platform services.\n\n* A controlled restart procedure was executed to ensure clean recovery of the impacted components.\n\n* Enhanced monitoring was maintained throughout the recovery process to validate platform stability and detect any recurrence.\n\n* Root cause analysis was completed and corrective actions were implemented to address the underlying infrastructure and resiliency gaps. \n\n\u200c\n\n**Future Preventative Measures**\n\n* Controlled Infrastructure Maintenance - Automatic upgrades have been disabled for the affected environment. Future infrastructure upgrades will be performed through planned maintenance activities with appropriate scheduling, oversight, and validation.\n\n* Increased Service Redundancy - The routing service has been scaled up to run multiple replicas across separate infrastructure nodes.\n\n* Improved Platform Resiliency - Platform resiliency controls have been reviewed and enhanced to better handle large-scale service restart events and reduce the risk of service discovery or routing bottlenecks during recovery operations.\n\n* Enhanced Monitoring and Early Detection - Additional monitoring and alerting have been implemented to provide earlier visibility into service degradation and infrastructure events that could impact platform availability",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-05-29T03:23:02.308-07:00",
"resolved_inferred": false,
"started_at": "2026-05-28T20:07:23.142-07:00",
"state": "postmortem",
"title": "Snow Atlas- APAC- Error 500 - internal server error",
"updated_at": "2026-06-16T01:41:35.706-07:00",
"url": "https://stspg.io/bxwnkmsz6xxy"
},
{
"body": "**Description:** Flexera One- IT Asset management - EU & APAC - PROD Reconcile failures\n\n**Timeframe:** May 26, 2026, 12:29 PM PDT to May 27, 2026, 06:32 AM PDT\n\n\u200c\n\n**Incident Summary**\n\n\u200c\n\nOn Wednesday, May 26, 2026, at 12:29 PM PDT, an issue was identified by our teams affecting Flexera One IT Asset Management customers in the EU and APAC production environments caused reconcile jobs to fail following a recent production release.\n\nFollowing a recent production release, a subset of Flexera One IT Asset Management customers in the EU and APAC production environments experienced failures during reconcile processing. The issue primarily affected the first reconcile run performed after the release. Customers encountering the issue received reconcile failures, resulting in delays to processing and an increase in support inquiries.\n\nUpon investigation, technical teams identified the likely source of the failures and confirmed that tenants which had already experienced an initial reconcile failure were expected to successfully complete subsequent reconcile runs. To prevent additional customers from being affected, a hotfix was developed, validated, and deployed to production. Following deployment of the hotfix, no additional reconcile failures related to this issue were observed. The platform remained under enhanced monitoring while teams validated system behavior and confirmed that reconcile processing was operating as expected. After an extended observation period with no further occurrences, the incident was declared resolved on May 27, 2026, at 06:32 AM PDT.\n\n\u200c\n\n**Root Cause**\n\n\u200c\n\nThe issue was caused by a data handling inconsistency introduced during a system update. Under specific conditions, this inconsistency resulted in unexpected behavior during the initial reconcile processing performed after the release, causing the reconcile job to fail.\n\nContributing Factors:\n\n* The issue only manifested during the first reconcile execution following the production release, making it difficult to identify during pre-release validation activities.\n\n* The affected processing path encountered conditions that were not fully represented in update validation scenarios.\n\n* Existing monitoring detected the reconcile failures but did not provide early indicators of the underlying processing condition before customer impact occurred.\n\n\u200c\n\n**Remediation Actions**\n\n\u200c\n\nThe following remediation steps were implemented to restore service functionality:\n\n* Technical teams investigated and identified the source of the reconcile failures.\n\n* Validation was performed to confirm that customers who had already encountered the issue could successfully complete subsequent reconcile runs.\n\n* A production hotfix was developed and deployed to prevent additional customers from experiencing the failure condition.\n\n* Enhanced monitoring was implemented during the recovery period to validate platform stability and reconcile processing behavior.\n\n* The incident remained under observation until teams confirmed no further customer impact was occurring.\n\n\u200c\n\n**Future Preventative Measures**\n\n\u200c\n\n* Enhanced Update Validation - Validation procedures for future system updates have been strengthened to improve coverage of data processing scenarios and compatibility checks that may occur during initial post-release operations.\n\n* Improved Monitoring and Detection - Additional monitoring and alerting opportunities have been identified and will be implemented to provide earlier visibility into reconcile processing anomalies and similar failure patterns.\n\n* Release Process Improvements - Lessons learned from this event have been incorporated into release planning and review processes to further reduce the risk of similar post-release processing issues and improve early detection should comparable symptoms arise in the future.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-05-27T06:33:08.174-07:00",
"resolved_inferred": false,
"started_at": "2026-05-27T04:04:19.417-07:00",
"state": "postmortem",
"title": "Flexera One- IT Asset management - EU & APAC - PROD Reconcile failures",
"updated_at": "2026-06-09T23:11:19.350-07:00",
"url": "https://stspg.io/01t4m2ynmmvt"
},
{
"body": "**Description:** Flexera One IT Asset Management \u2013 NAM \u2013 Beacon Upload and Download Issues\n\n**Timeframe:** May 22, 2026, 11:00 PM PDT \u2013 May 26, 2026, 11:41 PM PDT\n\n**Incident Summary**\n\nOn May 22, 2026, at approximately 11:00 PM PDT, an issue began affecting some Beacon activity for Flexera One IT Asset Management customers in the North America production environment.\n\nDuring the affected period, some customers experienced intermittent issues with inventory uploads and Beacon downloads. Reported symptoms included upload timeouts, inventory upload delays, increasing Beacon backlogs, and some download failures. Successful uploads and downloads continued to be observed during the incident, so the issue did not affect all Beacon activity.\n\nTechnical teams investigated the reported behavior by reviewing customer-reported failures, the Beacon request path, timeout patterns, backlog growth, and related backend processing errors. The investigation identified that some requests were not completing within the expected timeframe, which contributed to the intermittent behavior observed by affected customers.\n\nThe issue showed signs of backlog growth before the recent disaster recovery testing activity, with the impact increasing more significantly afterward. Based on the investigation, the disaster recovery activity contributed to the increased visibility and severity of the issue.\n\nAs part of the recovery effort, technical teams identified a connection handling issue in a backend service supporting Beacon processing. The affected service and related authentication components were restarted to restore normal processing behavior.\n\nFollowing these corrective actions, Beacon processing recovered, impacted backlogs began clearing, and affected upload and download functionality returned to expected operation. By May 26, 2026, at approximately 11:41 PM PDT, the issue had been resolved. Technical teams continued monitoring after recovery to confirm continued stability.\n\n**Root Cause**\n\nThe incident was caused by a connection handling issue in a backend service supporting Beacon processing. This contributed to intermittent upload timeouts, inventory upload delays, Beacon backlog growth, and some Beacon download failures for affected customers.\n\n**Remediation Actions**\n\nThe following actions were taken during the incident response:\n\n1. **Incident Response Initiated:** Technical teams began investigating customer reports of Beacon upload and download issues in the North America production environment.\n2. **Impact Assessed:** Teams reviewed customer-reported symptoms and confirmed the issue was intermittent, with successful Beacon activity still being observed.\n3. **Technical Investigation Performed:** Teams reviewed the affected Beacon request path, timeout behavior, backlog growth, and related backend errors to isolate the cause.\n4. **Corrective Action Applied:** The affected backend service and related authentication components were restarted to restore normal processing behavior.\n5. **Recovery Validated:** Teams confirmed that Beacon upload and download functionality had recovered and that impacted backlogs were clearing.\n6. **Post-Recovery Monitoring Continued:** Teams continued monitoring after restoration to confirm continued stability.\n\n**Future Preventative Measures**\n\nThis incident highlighted the importance of continued resilience for services supporting Beacon upload and download workflows, particularly when processing delays or backlog growth occur.\n\nBased on the investigation, the following follow-up activities have been identified:\n\n1. **Connection Handling Review:** Review connection handling behavior for the backend service supporting Beacon processing to reduce the likelihood of similar issues recurring.\n2. **Processing Resilience Review:** Evaluate opportunities to improve resilience when Beacon-related requests take longer than expected to complete.\n3. **Beacon Workflow Validation Review:** Review validation steps for Beacon upload and download workflows following major operational activities, including disaster recovery testing.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-05-27T00:33:15.311-07:00",
"resolved_inferred": false,
"started_at": "2026-05-26T14:49:13.824-07:00",
"state": "postmortem",
"title": "Flexera One IT Asset Management \u2013 NAM \u2013 Beacon Upload and Download Issues",
"updated_at": "2026-06-10T23:47:24.576-07:00",
"url": "https://stspg.io/nssn6ldvfy15"
},
{
"body": "**Description:** Flexera One - IT Visibility - NA - Delayed Inventory Data Updates\n\n**Timeframe:** May 15, 2026, 8:40 AM PDT - May 25, 2026, 9:08 PM PDT\n\n**Incident Summary**\n\nOn May 15, 2026, at approximately 8:40 AM PDT, we received customer reports of delayed IT Visibility inventory data updates in the North America production environment.\n\nThe reports were received after a previous backlog-related incident had been closed. Some customers continued to observe inventory imports displaying a timeout status, while others reported delays in updated inventory data becoming available within the application.\n\nTechnical teams investigated the reported behavior to determine whether the issue was related to residual processing activity from the previous incident, a recurrence of earlier processing delays, or a separate import status or processing concern.\n\nDuring the investigation, technical teams identified processing-related conditions that contributed to delayed completion for some inventory imports. This included processing slowness and errors in supporting services that could prevent some inventory processing activity from completing within expected timeframes.\n\nCorrective actions were applied to improve processing behavior, and affected processing activity was reviewed and reprocessed where needed. Technical teams continued validating reported customer examples and monitoring new inventory processing activity to confirm recovery.\n\nBy May 28, 2026, at approximately 9:08 PM PDT, broader platform activity and processing had returned to expected levels, and the incident was considered resolved.\n\n**Root Cause**\n\nThe incident was caused by multiple contributing factors affecting the timely processing and availability of updated IT Visibility inventory data.\n\nSome inventory imports experienced delayed completion due to processing slowness and errors in supporting services. As a result, updated inventory data was delayed in becoming available within IT Visibility for some customers.\n\nIn addition, some imports displayed a timeout status when processing took longer than expected. In some cases, this timeout status did not indicate that processing had stopped, as processing could continue in the background and complete later.\n\n**Remediation Actions**\n\nThe following actions were taken during the incident response:\n\n1. **Incident Response Initiated:** Technical teams began investigating after customer reports were received regarding timeout status behavior and delayed inventory data updates.\n2. **Customer Reports Reviewed:** Reported customer examples and related support cases were reviewed to validate the scope and current behavior.\n3. **Processing Activity Validated:** Technical teams reviewed inventory processing activity to determine whether timeout statuses reflected active processing delays or status behavior after processing had continued in the background.\n4. **Processing Fix Applied:** A corrective fix was implemented for an identified issue that could contribute to timeout behavior, and affected processing activity was reprocessed.\n5. **Supporting Service Errors Investigated:** Technical teams identified errors in supporting services that were contributing to delayed processing completion and engaged the appropriate teams for review and recovery.\n6. **Processing Recovery Monitored:** Teams monitored affected processing activity and confirmed that delayed processing continued improving as recovery progressed.\n7. **Customer Examples Validated:** Technical teams continued validating reported customer examples to confirm whether updated inventory data was progressing as expected.\n8. **Post-Recovery Monitoring Performed:** Teams continued monitoring platform activity and inventory processing behavior before declaring the incident resolved.\n\n**Future Preventative Measures**\n\nThis incident highlighted opportunities to improve visibility, monitoring, and recovery handling for IT Visibility inventory processing.\n\nThe following follow-up areas are being reviewed:\n\n1. **Inventory Processing Reliability:** Technical teams are reviewing improvements to strengthen inventory processing reliability and reduce the likelihood of similar delays.\n2. **Monitoring and Observability:** Additional monitoring improvements are being evaluated to provide earlier visibility into processing delays, timeout patterns, backlog, and lack of processing progress.\n3. **Import Status Clarity:** Teams are reviewing import status behavior to help reduce confusion when processing takes longer than expected but may still continue in the background.\n4. **Operational Readiness:** Teams are reviewing operational improvements to support faster investigation, clearer internal visibility, and more consistent recovery handling for similar issues in the future.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-05-25T21:39:28.696-07:00",
"resolved_inferred": false,
"started_at": "2026-05-15T10:11:50.682-07:00",
"state": "postmortem",
"title": "Flexera One - IT Visibility - NA - Delayed Inventory Data Updates",
"updated_at": "2026-08-09T12:29:38.699-07:00",
"url": "https://stspg.io/tzj2nw2xmrgy"
},
{
"body": "**Description:** Spot Elastigroup \u2013 AWS \u2013 Page Load Delays and Creation Issues\n\n**Timeframe:** May 13, 2026, 7:29 AM PDT to May 13, 2026, 9:38 AM PDT\n\n\u200c\n\n**Incident Summary**\n\n\u200c\n\nOn Wednesday, our teams identified an issue affecting AWS Elastigroup functionality that impacted Spot Elastigroup for AWS customers, resulting in delays when loading the Elastigroup page and preventing some customers from creating new Elastigroups. The impact was limited to AWS Elastigroup functionality, and validation confirmed that other services were not affected.\n\nTechnical teams immediately initiated an investigation, reviewing recent platform changes and service behavior. During the investigation, they observed instability in services that correlated with the customer-facing symptoms. In response, teams executed mitigation actions, including reverting a recently introduced change while concurrently validating system behavior and overall service health.\n\nFollowing the rollback, validation confirmed that Elastigroup creation functionality had been successfully restored. Additional performance validation and monitoring confirmed that service behavior had returned to expected levels, with no further customer impact observed. Full functionality was restored by 9:38 AM PDT.\n\n\u200c\n\n**Root Cause**\n\n\u200c\n\nThe issue was caused by a recently introduced change that resulted in instability within the AWS Elastigroup processing path. This behavior affected service interactions required for Elastigroup page loading and new Elastigroup creation requests, leading to increased latency and request failures for impacted customers.\n\nContributing Factors:\n\n* The change unexpectedly affected a broad AWS workload scope, increasing the overall impact radius.\n* Existing monitoring did not provide early detection for repeated pod restart behavior.\n* Existing alerting mechanisms did not immediately identify the developing service degradation.\n\n\u200c\n\n**Remediation Actions**\n\n\u200c\n\nThe following remediation actions were taken to restore service:\n\n* Investigated customer-reported Elastigroup page loading delays and creation failures.\n* Reviewed system behavior and service restart activity impacting the AWS processing path.\n* Reverted the recently introduced change.\n* Performed validation of Elastigroup creation workflows through automated testing.\n* Conducted additional performance validation and environment monitoring to confirm stability.\n\n\u200c\n\n**Future Preventative Measures**\n\n\u200c\n\n1. Improved Deployment Safeguards - Strengthen deployment controls through stricter practices and introduce automatic rollback mechanisms when latency or error thresholds are exceeded.\n2. Defined Rollback Strategy \u2013 Review and standardize rollback playbook with clearly defined triggers, including error-rate and latency thresholds, to enable faster mitigation during service degradation events.\n3. Enhanced Service Health Monitoring - Implement additional alerting and monitoring for continuous pod restart activity to improve early detection and reduce response times for emerging issues.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-05-13T10:02:51.973-07:00",
"resolved_inferred": false,
"started_at": "2026-05-13T07:49:21.000-07:00",
"state": "postmortem",
"title": "Spot Elastigroup \u2013 AWS \u2013 Page Load Delays and Creation Issues",
"updated_at": "2026-05-26T10:55:27.884-07:00",
"url": "https://stspg.io/6vvgmndj0gfp"
},
{
"body": "**Description:** Flexera One- IT Visibility- US Prod - Degraded performance\n\n**Timeframe:** May 10, 2026, 8:12 PM PDT to May 14, 2026, 07:25 PM PDT\n\n\u200c\n\n**Incident Summary**\n\n\u200c\n\nOn Sunday,  May 10 , 2026, at 8:12 PM PDT , our teams detected an issue impacting IT Visibility \\(ITV\\) services in the US region where the affected customers experienced delays in normalized inventory processing and delivery to ITV UIs and APIs.\n\nThe issue originated within the normalization persistence layer, where the service encountered repeated failures while initializing streaming clients used for communication. During initialization, the service generated a large number of API requests in a short period of time, which exceeded account throttling limits. As a result, the service repeatedly restarted and was unable to consistently process and persist normalized inventory data.\n\nOur technical teams identified the contributing factors, implemented mitigation measures, and deployed code improvements designed to reduce API request spikes and improve service resiliency during startup and recovery conditions.\n\nFollowing deployment of the fixes, services stabilized and backlog processing was initiated in a controlled manner to avoid downstream system impact. Recovery progressed steadily, and all backlog processing was successfully completed by May 14, 2026, at 07:25 PM PDT.\n\n\u200c\n\n**Root Cause**\n\n\u200c\n\n* The incident was caused by excessive API requests generated during streaming client initialization within the normalization writer service. The request volume exceeded throttling limits, preventing successful initialization and causing the service to repeatedly restart.\n\nContributing Factors:\n\n* The production US environment had significantly scaled up,  increasing the number of streaming clients initialized during service startup.\n* Separate streaming clients were created per organization across multiple collections, resulting in unexpectedly high API calls during each pod restart.\n* Failure handling logic caused the service to terminate immediately instead of retrying gracefully, amplifying restart frequency and additional request spikes.\n\n\u200c\n\n**Remediation Actions**\n\n\u200c\n\nThe following remediation activities were completed to restore service stability:\n\n* Implemented staggered streaming client initialization to reduce API request spikes during service startup.\n* Added retry logic with exponential backoff around streaming client creation and authentication requests.\n* Replaced failure handling with graceful recovery and retry mechanisms.\n* Stabilized normalization services and resumed backlog processing in a controlled manner.\n* Closely monitored recovery activities to ensure downstream platform stability during backlog replay.\n\n\u200c\n\n**Future Preventative Measures**\n\n* Introduce rate limiting controls for external dependency initialization workflows.\n* Enhance resiliency standards for service startup and dependency authentication handling.\n* Improve observability and alerting around throttling conditions.\n* Review scalability assumptions and startup behavior for high-scale growth scenarios.\n* Conduct additional resiliency testing focused on restart  and dependency throttling conditions.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-05-14T21:42:12.543-07:00",
"resolved_inferred": false,
"started_at": "2026-05-10T20:31:38.321-07:00",
"state": "postmortem",
"title": "Flexera One- IT Visibility- US Prod - Degraded performance",
"updated_at": "2026-05-29T03:50:59.958-07:00",
"url": "https://stspg.io/j6y9gkp8fdlw"
},
{
"body": "**Description:** Flexera One \u2013 Cloud Cost Optimization \u2013 NAM \u2013 Intermittent Service Degradation and Cost Data Processing Delays\n\n**Timeframe:** May 7, 2026, 11:49 PM PDT \u2013 May 8, 2026, 8:04 PM PDT\n\n**Incident Summary**\n\nOn May 7, 2026, at approximately 11:49 PM PDT, an issue was identified affecting Flexera One Cloud Cost Optimization customers in the North America region.\n\nDuring the affected period, some customers may have experienced intermittent errors when accessing Cloud Cost Optimization functionality through the Flexera One user interface. Some customers may also have experienced delays in cost data processing.\n\nThe issue was identified through internal technical observations, including increased Cloud Cost Optimization cost API request failures and data processing delays, rather than through widespread customer reports.\n\nThe incident was initially investigated as a broader Flexera One service degradation because intermittent loading behavior was being reviewed across multiple areas of the platform during the same period. As a precaution, additional components were included in the initial customer communications while technical teams validated the affected areas and determined whether the reported behaviors were related to the same underlying service provider disruption.\n\nFurther investigation confirmed that the service provider disruption was associated with the Cloud Cost Optimization impact described in this report. Other service behavior reviewed during the investigation was determined to be separate from the service provider disruption and was handled through separate investigation and tracking.\n\nTechnical teams continued monitoring Cloud Cost Optimization service behavior while the upstream disruption was being addressed. After the service provider disruption was resolved, technical teams validated Cloud Cost Optimization functionality and confirmed that services had recovered. By May 8, 2026, at approximately 8:04 PM PDT, the incident was considered resolved.\n\n**Root Cause**\n\nThe incident was caused by a disruption affecting an upstream service provider. As a result, Cloud Cost Optimization experienced intermittent cost API request failures and data processing delays in the North America region.\n\n**Remediation Actions**\n\nThe following actions were taken during the incident response:\n\n1. **Incident Investigation Initiated:** Technical teams began investigating intermittent errors and cost data processing delays affecting Cloud Cost Optimization customers in the North America region.\n2. **Internal Observations Reviewed:** Technical teams reviewed internal observations, including increased Cloud Cost Optimization cost API request failures and data processing delays.\n3. **Initial Scope Reviewed:** Because intermittent loading behavior was being reviewed across multiple areas of Flexera One during the same period, the incident was initially assessed as a broader platform service degradation.\n4. **Customer Communications Posted:** Additional components were included in the initial customer communications as a precaution while technical teams validated the confirmed scope of impact.\n5. **Service Provider Disruption Identified:** Technical teams determined that the confirmed Cloud Cost Optimization impact was associated with an upstream service provider disruption.\n6. **Impact Assessment Performed:** Technical teams assessed Cloud Cost Optimization API request behavior, Flexera One user interface behavior, and cost data processing delays.\n7. **Additional Service Behavior Reviewed:** Additional service behavior reported during the same period was reviewed to determine whether it was related to the same service provider disruption. This behavior was determined to be separate from the service provider disruption and was handled through separate investigation and tracking.\n8. **Monitoring Continued:** Technical teams continued monitoring Cloud Cost Optimization service behavior while the upstream service provider disruption was being addressed.\n9. **Recovery Validated:** After the upstream service provider disruption was resolved, technical teams validated Cloud Cost Optimization functionality and confirmed recovery before resolving the incident.\n\n**Future Preventative Measures**\n\nThis incident highlighted the importance of continued validation when multiple service behaviors are reported during a broader upstream service provider disruption.\n\nBased on the investigation, the following follow-up activity has been identified:\n\n1. **Incident Triage and Component Validation Review:** We will review how related service reports are assessed during broad upstream service provider disruptions to ensure affected components are validated and updated as additional information becomes available. This includes continuing to refine triage practices so customer communications remain aligned with the confirmed scope of impact.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-05-09T02:47:15.005-07:00",
"resolved_inferred": false,
"started_at": "2026-05-08T15:44:35.760-07:00",
"state": "postmortem",
"title": "Flexera One \u2013 NAM \u2013 Intermittent Service Degradation",
"updated_at": "2026-05-25T22:39:29.745-07:00",
"url": "https://stspg.io/d6ytccjqgkvz"
},
{
"body": "**Description:** Flexera- Spot- US east- 1a- Database connection and low request issues\n\n**Timeframe:** May 7, 2026, 4:25 PM PST to May 8, 2026, 7:04 PM PST\n\n\u200c\n\n**Incident Summary**\n\n\u200c\n\nOn May 7, 2026, Flexera identified a service degradation impacting Spot services in the US East 1a region. During the impact duration, customers experienced difficulties launching instances. In addition, some backend services experienced reduced request handling capacity and intermittent database connectivity issues, resulting in degraded provisioning behavior and delayed scaling operations.\n\nInitial investigation determined that the issue coincided with a broader disruption affecting an external cloud infrastructure provider. The external disruption introduced elevated latency and intermittent service instability across dependent infrastructure components, which contributed to degraded provisioning and backend service performance within the Spot platform.\n\nDuring the investigation, engineering teams also identified a database issue caused by the service provider outage that posed a potential performance risk under degraded infrastructure conditions. The issue was mitigated early in the response process, and system stability improved following remediation actions.\n\nThroughout the incident, technical teams continuously monitored service health, validated provisioning behavior, and assessed mitigation options to minimize customer impact while the external provider disruption remained active. Over several hours following mitigation, Spot services remained stable with no recurrence of customer-facing issues observed.\n\nAfter an extended monitoring period confirmed continued stability, the incident was considered resolved following confirmation that the underlying external provider disruption had been fully remediated on May 8, 2026, at 7:04 PM PST.\n\n\u200c\n\n**Root Cause**\n\n\u200c\n\nThe issue was primarily caused by a disruption affecting an external cloud infrastructure provider supporting services in the US East 1a region. The disruption resulted in increased latency, intermittent connectivity issues, and degraded infrastructure performance, which impacted Spot instance provisioning operations and related backend services.\n\nContributing Factors:\n\n* Elevated latency and intermittent failures across dependent infrastructure services increased provisioning delays and request instability. \n* A database performance issue, triggered by the underlying service provider disruption, introduced additional load under degraded infrastructure conditions and increased the risk of intermittent service instability. \n* The prolonged nature of the external provider disruption extended the duration of customer impact and required continuous monitoring and mitigation activities.\n\n\u200c\n\n**Remediation Actions**\n\n\u200c\n\nThe following remediation steps were implemented to restore service functionality:\n\n* Investigated and monitored infrastructure and provisioning service behavior across impacted Spot components. \n* Identified and mitigated a database performance degradation contributing to elevated performance risk. \n* Validated provisioning stability and backend service recovery following mitigation efforts. \n* Performed extended monitoring to ensure sustained service stability and confirm the absence of recurring customer impact. \n\n\u200c\n\n**Future Preventative Measures**\n\n\u200c\n\n* Alerting Reliability Improvements: Improve alert delivery validation and monitoring to ensure critical database alerts are consistently propagated across all operational notification channels, including collaboration and incident management platforms.\n* Regional Resilience Evaluation: Evaluate additional regional resilience and failover capabilities to reduce the impact scope of regional infrastructure disruptions. Any future implementation will be subject to cost-impact analysis and internal approval processes.\n* Third-Party Provider Escalation Process: Enhance operational procedures for third-party infrastructure incidents by establishing earlier escalation and engagement processes with external cloud service providers during active service disruptions.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-05-08T22:23:32.602-07:00",
"resolved_inferred": false,
"started_at": "2026-05-07T20:00:40.000-07:00",
"state": "postmortem",
"title": "Flexera- Spot- US east- 1- Service Degradation (Instance Launch Failures)",
"updated_at": "2026-05-13T03:23:14.379-07:00",
"url": "https://stspg.io/33hp3s8xdnst"
},
{
"body": "**Description:** Flexera One \u2013 IT Asset Management \u2013 North America \u2013 Beacon Connectivity Issues\n\n**Timeframe:** May 6, 2026, 9:33 PM PDT \u2013 May 7, 2026, 11:35 AM PDT\n\n## Incident Summary\n\nOn Wednesday, May 6, 2026, at approximately 9:33 PM PDT, a scheduled backend infrastructure change was performed in the North America production environment. Following this change, an issue was introduced that affected Beacon connectivity for Flexera One IT Asset Management customers in the North America region.\n\nDuring the affected period, customers experienced failures when attempting to upload or download data through Beacon. The issue affected Beacon API and configuration service connectivity, preventing successful uploads and downloads across the US production environment.\n\nThe following morning, technical teams began investigating the issue after customer reports were received. During the investigation, it was identified that a configuration introduced during the infrastructure change was directing Beacon traffic to an incorrect backend path, preventing successful connectivity.\n\nOnce the configuration was corrected, Beacon uploads and downloads began processing successfully. Multiple customers confirmed successful Beacon connectivity following the fix. By Thursday, May 7, 2026, at approximately 11:35 AM PDT, service was confirmed restored and the incident was considered resolved. Affected Beacon uploads would have automatically retried following service restoration. Technical teams continued monitoring following restoration to ensure recovery progressed as expected.\n\n## Root Cause\n\nThe incident was caused by a backend infrastructure change that introduced a configuration mismatch in the North America production environment. As a result, Beacon connectivity requests were directed to an incorrect backend path, causing all Beacon uploads and downloads to fail during the affected period. Once the configuration was corrected to point to the appropriate path, connectivity was restored.\n\n## Remediation Actions\n\nThe following actions were taken during the incident response:\n\n1. **Configuration Corrected:** The affected configuration was updated to restore Beacon connectivity to the correct backend path.\n2. **Beacon Connectivity Validated:** Technical teams confirmed that Beacon uploads and downloads were processing successfully following the correction.\n3. **Customer Confirmation Received:** Multiple customers confirmed successful Beacon connectivity.\n4. **Post-Restoration Monitoring Performed:** Recovery activity was monitored after restoration to ensure services continued operating as expected.\n\n## Future Preventative Measures\n\nThe following follow-up actions were identified during the incident:\n\n1. **Cross-Team Change Coordination:** We are reviewing coordination procedures between teams involved in backend infrastructure changes to ensure that dependent services are validated before and after changes are applied.\n2. **Remaining Region Updates:** Configuration updates for other regions are being coordinated as part of upcoming scheduled releases to ensure consistency and prevent similar occurrences.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-05-07T12:38:18.010-07:00",
"resolved_inferred": false,
"started_at": "2026-05-07T09:49:00.000-07:00",
"state": "postmortem",
"title": "Flexera One IT Asset Management \u2013 NAM \u2013 Beacon Connectivity Issues",
"updated_at": "2026-05-21T11:31:40.465-07:00",
"url": "https://stspg.io/t3hsypf1tnbj"
},
{
"body": "**Description:** Flexera One SaaS Manager \u2013 NAM \u2013 Application Access Disruption\n\n**Timeframe:** April 29, 2026, 3:00 AM PDT \u2013 April 29, 2026, 7:30 AM PDT\n\n## Incident Summary\n\nOn April 29, 2026, at approximately 3:00 AM PDT, customers began experiencing access issues with Flexera One SaaS Manager in the North America production environment.\n\nDuring this period, affected customers were unable to access SaaS Manager. The issue was related to an internal SaaS Manager configuration update process that did not complete within the expected timeframe for affected organizations.\n\nTechnical teams investigated and identified that a high volume of internal SaaS Manager configuration updates had been triggered for customer organizations within a short period of time. These updates did not complete quickly enough, causing affected organizations to remain in an updating state and preventing access to SaaS Manager.\n\nAs part of the recovery effort, technical teams increased processing capacity for the affected update process and worked through the pending configuration updates. Once the pending updates were completed, affected organizations were no longer in an updating state and access to SaaS Manager was restored.\n\nBy April 29, 2026, at approximately 7:30 AM PDT, access had been restored for affected customers. Following validation that SaaS Manager access had returned to expected operation, the incident was considered resolved.\n\n## Root Cause\n\nThe incident was caused by a high volume of internal SaaS Manager configuration updates for customer organizations that did not complete within the expected timeframe. While these updates remained pending, affected customer organizations were left in an updating state, which prevented access to SaaS Manager during the incident window.\n\n## Remediation Actions\n\nThe following actions were taken during the incident response:\n\n* **Incident Investigation Initiated:** Technical teams investigated reports of SaaS Manager access issues affecting customers in the North America region.\n* **Impact Scope Reviewed:** The issue was confirmed to affect customers in the NAM region.\n* **Configuration Update Processing Reviewed:** Technical teams identified that a high volume of internal SaaS Manager configuration updates had been triggered and were not completing within the expected timeframe.\n* **Processing Capacity Increased:** Technical teams increased processing capacity for the affected update process to help complete the pending configuration updates.\n* **Pending Updates Completed:** The pending configuration updates were processed for affected organizations.\n* **Access Restoration Completed:** Once the pending updates were completed, affected organizations were no longer in an updating state and access to SaaS Manager was restored.\n* **Service Recovery Validated:** Technical teams confirmed that access to SaaS Manager had returned to expected operation for affected organizations.\n\n## Future Preventative Measures\n\nThis incident highlighted the importance of improving resilience, capacity handling, and monitoring for high-volume internal configuration update activity that can affect customer access.\n\nBased on the investigation, the following follow-up activities are being reviewed:\n\n* **Capacity Handling Review:** Review processing behavior for high-volume configuration update activity to reduce the likelihood of pending updates affecting customer access.\n* **Update Process Resilience:** Evaluate improvements to how SaaS Manager configuration updates are processed so that customer access is not unnecessarily affected if update activity takes longer than expected.\n* **Monitoring and Alerting Improvements:** Review monitoring opportunities to identify large volumes of pending or delayed configuration updates earlier.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "critical",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-04-29T11:33:50.000-07:00",
"resolved_inferred": false,
"started_at": "2026-04-29T11:33:50.000-07:00",
"state": "postmortem",
"title": "First and Final Notification: Flexera One SaaS Manager \u2013 NAM \u2013 Application Access Disruption",
"updated_at": "2026-06-07T18:19:29.540-07:00",
"url": "https://stspg.io/c91hwx071r1n"
},
{
"body": "**Description:** Flexera One \u2013 IT Asset Management \u2013 NA \u2013 Reconciliation Processing Delays\n\n**Timeframe:** April 28, 2026, 6:10 AM PDT \u2013 April 28, 2026, 10:45 AM PDT  \n\n**Incident Summary**\n\nOn April 28, 2026, at approximately 6:10 AM PDT, an issue was identified affecting reconciliation processing within the Flexera One IT Asset Management service in the North America production environment.  \n  \nDuring this period, the application remained accessible. However, affected customers may have experienced delays in reconciliation completion and related writer job processing. Investigation identified that some writer jobs were blocked while processing reconciliation results, with repeated failures observed during result streaming.  \n  \nTechnical teams began investigating immediately and reviewed logs associated with the affected processing flow. The investigation identified repeated attempts to retrieve reconciliation results after streaming errors occurred, resulting in a large volume of repeated errors. Several writer jobs were also observed running longer than expected.  \n  \nAs part of the recovery effort, a service restart or redeploy was attempted but did not resolve the issue. Technical teams then identified that a recent service release occurred around the time processing load began increasing. The service was reverted to a previous version as part of the mitigation effort, and further deployment activity was held while the issue was reviewed.  \n  \nAdditional recovery actions were taken for writer jobs that were already stuck in an unrecoverable state. These jobs were cleared where needed, while newly started writer jobs were monitored to confirm that they were processing successfully. By April 28, 2026, at approximately 10:45 AM PDT, current writer job processing was confirmed to be completing successfully, and the incident was considered resolved. Any longer-running jobs not related to the affected processing stage continued to be monitored separately.\n\n  \n**Root Cause**\n\nThe incident was caused by writer jobs entering a blocked state while processing reconciliation in the North America production environment. This resulted in repeated failures during result streaming and delayed completion of reconciliation-related writer job processing for affected customers.\n\nThe issue occurred around the time of a service release, and technical teams identified that a deployment during active reconciliation processing may cause certain writer jobs to enter an unrecoverable error state. When this occurred, retry behavior from the IT Asset Management processing flow continued until the affected jobs timed out or were manually cleared.\n\n  \n**Remediation Actions**\n\nThe following actions were taken during the incident response:\n\n1. Incident Investigation Initiated: Technical teams began investigating delays affecting reconciliation completion and related writer job processing in the North America region. \n2. Log Review Completed: Logs were reviewed and showed repeated failures during reconciliation result streaming, including repeated attempts to retrieve results after errors occurred. \n3. Blocked Writer Jobs Identified: Technical teams identified that some writer jobs were blocked during reconciliation result processing and had been running longer than expected. \n4. Service Restart/Redeploy Attempted: A restart or redeploy of the affected service was attempted but did not resolve the issue, as affected writer jobs had already entered a blocked processing state.\n5. Recent Service Release Reverted: A recent service release that occurred around the time processing load increased was reverted to a previous version as part of the mitigation effort. \n6. Further Deployment Activity Held: Additional deployment activity was paused while technical teams reviewed the behavior and assessed the safest recovery approach. \n7. Stuck Writer Jobs Cleared: Writer jobs that were confirmed to be stuck in an unrecoverable error state were cleared where needed. \n8. Processing Recovery Validated: Newly started and current writer jobs were monitored and confirmed to be completing successfully before the incident was resolved.\n\n  \n**Future Preventative Measures**\n\nThis incident highlighted the importance of improving resilience in reconciliation processing when service changes occur during active writer job processing.  \nBased on the investigation, the following follow-up activities are being pursued:\n\n1. Deployment Strategy Review: Review deployment approaches for the affected service to reduce the likelihood of disrupting active reconciliation processing.\n2. Retry Behavior Improvements: Review retry behavior for reconciliation-related writer jobs to reduce repeated processing attempts when a job reaches an unrecoverable error condition.\n3. Failure Handling Improvements: Evaluate code-side improvements to better detect specific failure patterns and allow affected processing to fail or recover more cleanly.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-04-28T10:50:21.661-07:00",
"resolved_inferred": false,
"started_at": "2026-04-28T08:31:03.628-07:00",
"state": "postmortem",
"title": "Flexera One - IT Asset Management - NA - Reconciliation Processing Delays",
"updated_at": "2026-05-11T22:43:29.621-07:00",
"url": "https://stspg.io/dqbwdpfw2fzb"
},
{
"body": "**Description:** Snow Software - Australia - SAM Core \\(Snow Atlas\\) Errors\n\n**Timeframe:** April 21, 2026, 1:35 AM PST to April 21, 2026, 2:05 AM PST\n\n\u200c\n\n**Incident Summary**\n\n\u200c\n\nOn Tuesday, April 21, 2026, 1:35 AM PST ,customers in the Australia region began experiencing HTTP 500 errors when accessing specific SAM Core functionality, particularly within the Application and License detail views. Although the Snow Atlas portal itself remained accessible, certain features failed to load, disrupting normal user operations. Initial investigation indicated a potential issue in the messaging system, supported by logs from the API service that reported a routing error.\n\nAs a mitigation step, the route application in the Australia region was restarted at 1:38 AM PST, which led to a temporary improvement in system behaviour. The issue, however, reoccurred shortly thereafter, confirming that the restart did not resolve the underlying problem.\n\nSubsequent investigation linked the incident to a recent change associated with regional migration activities between Australia Southeast and Australia East. The team determined this change was a likely contributing factor and initiated a rollback of the prior day\u2019s deployment.\n\nFollowing the rollback, system stability improved, and by 2:05 AM PST, validation with multiple previously impacted customers confirmed that the affected pages were once again functioning as expected. Full service was restored, and the incident was considered resolved.\n\n\u200c\n\n**Root Cause**\n\n\u200c\n\nThe issue was caused by an unintended side effect of recent changes introduced during the Australia Southeast to Australia East migration, which resulted in a routing failure within the application layer.\n\nContributing Factors:\n\n* A recent deployment related to migration cleanup activities introduced instability in routing behavior.\n* A fatal routing error in the API service disrupted request handling.\n* Potential interaction with the messaging system contributed to service degradation \\(under investigation\\).\n* The issue was not immediately reproducible during initial validation after the change, delaying detection.\n\n\u200c\n\n**Remediation Actions**\n\n\u200c\n\nThe following remediation steps were implemented to restore service functionality:\n\n* Restarted the routing service to attempt initial recovery.\n* Conducted detailed log analysis to identify routing failures.\n* Rolled back the prior day\u2019s deployment associated with migration changes.\n* Validated service recovery with affected customers.\n\n\u200c\n\n**Future Preventative Measures**\n\n\u200c\n\n* Enhance validation and monitoring for regional migration-related changes.\n* Implement additional safeguards and automated checks for routing and messaging dependencies.\n* Introduce stricter post-deployment verification processes to detect delayed failures.\n* Conduct a detailed post-mortem to identify any additional contributing conditions and ensure long-term stability.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-04-21T02:23:10.872-07:00",
"resolved_inferred": false,
"started_at": "2026-04-21T02:01:34.600-07:00",
"state": "postmortem",
"title": "Snow Software - Australia - Specific pages not loading",
"updated_at": "2026-05-04T23:29:54.124-07:00",
"url": "https://stspg.io/thgn2fy2jplr"
},
{
"body": "**Description:** Flexera One - IT Asset Management - NAM - Intermittent 500 error\n\n**Timeframe:** April 13, 2026, 1:46 AM PST to Apr 13, 2026, 2:27 AM PST\n\n\u200c\n\n**Incident Summary**\n\n\u200c\n\nOn Monday, April 13, 2026, at 1:46 AM PST , our teams identified an issue affecting beacon communications for a subset of customers in the NAM region. During the incident window, customers may have experienced intermittent failures when communicating with the Flexera One platform through beacon services, including inventory uploads, downloads, and data import operations. Affected requests returned HTTP 500 \\(Internal Server Error\\) responses. Our teams also identified that some customers might have experienced the issue intermittently upto a few days before issue detection.\n\nThe issue was traced to disk space exhaustion on shared storage utilized by the beacon  processing infrastructure. Once the available disk capacity was exhausted, the affected services were unable to successfully process beacon-related requests, resulting in intermittent failures.\n\nOur technical teams quickly identified the storage exhaustion condition and restored service by expanding disk capacity on the affected infrastructure. \n\nFollowing remediation, beacon communications returned to normal operation and the environment was closely monitored to ensure continued stability.\n\n\u200c\n\n**Root Cause**\n\n\u200c\n\nThe incident was caused by disk space exhaustion on shared storage used by the beacon  processing component. Once the storage became full, beacon processing requests began failing and returning HTTP 500 responses.\n\nContributing Factors\n\n* Existing cleanup mechanisms did not fully remove all temporary or residual data generated during processing activities.\n* Under certain application failure or unexpected processing scenarios, files were retained longer than intended, contributing to accelerated disk utilization growth.\n* Storage utilization alerts did not trigger as expected due to a recent update introduced on the alerting mechanism.   \n\n**Remediation Actions**\n\n\u200c\n\nThe following remediation actions were completed during the incident response:\n\n* Expanded disk capacity on the affected shared storage infrastructure.\n* Removed the problematic configuration on the alert notification mechanism to restore alert delivery functionality.\n* Validated restoration of beacon communications and successful processing of requests.\n* Monitored the environment following recovery to confirm service stability.\n\n\u200c\n\n**Future Preventative Measures**\n\n\u200c\n\n* To reduce the likelihood of recurrence, the following preventative measures are being implemented:\n* Enhance cleanup and retention mechanisms to ensure all relevant temporary and residual processing files are automatically removed.\n* Improve storage utilization monitoring and alerting coverage across shared  processing infrastructure.\n* Validate alert delivery configurations after infrastructure changes.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-15T12:28:18Z",
"resolved_at": "2026-04-14T05:15:54.000-07:00",
"resolved_inferred": false,
"started_at": "2026-04-14T05:15:54.000-07:00",
"state": "postmortem",
"title": "Flexera One - IT Asset Management - NAM - Intermittent 500 errors",
"updated_at": "2026-05-12T02:20:51.352-07:00",
"url": "https://stspg.io/w1qq8dls6mz4"
},
{
"body": "**Description:** Flexera One \u2013 IT Asset Management \u2013 EU \u2013 Inventory Upload Failures\n\n**Timeframe:** April 7, 2026, 10:53 AM PDT \u2013 April 8, 2026, 7:45 AM PDT  \n\n**Incident Summary**\n\nOn April 7, 2026, at approximately 10:53 AM PDT, an issue began affecting inventory data uploads for Flexera One IT Asset Management in the EU production environment. During this period, customers in the EU region experienced failures when attempting to upload inventory data, which also resulted in delays in downstream data processing.  \n  \nTechnical teams began investigating the issue after reports were received of upload failures affecting the EU region. During the investigation, it was confirmed that the upload endpoint was reachable and responding, while customer upload attempts were still failing. Further analysis continued to identify the source of the disruption and determine why uploads were not being processed successfully.  \n  \nThe issue was subsequently identified as a required service in the EU production environment being unexpectedly stopped. As a result, upload requests were not processed successfully during the affected period, which caused authentication failures and prevented data from being received for processing. Once the service was restarted, uploads began processing successfully again and recovery activity was monitored.  \n  \nBy April 8, 2026, at approximately 7:45 AM PDT, successful uploads had resumed and the incident was considered resolved. It was confirmed during the incident that there was no data loss, as affected uploads would automatically re-upload after service restoration. Technical teams continued monitoring following restoration to ensure recovery progressed as expected.   \n\n**Root Cause**\n\nThe incident was caused by a required service in the EU production environment being unexpectedly stopped. This interruption prevented inventory upload requests from being processed successfully, resulting in authentication failures and no data being received for processing during the affected period. Further analysis is underway to determine the cause of the unexpected service interruption.  \n\n**Remediation Actions**\n\nThe following actions were taken during the incident response:\n\n1. Incident Investigation Initiated: Technical teams began investigating reports of inventory upload failures affecting the EU region. \n2. Endpoint Behavior Reviewed: Technical teams confirmed that the upload endpoint was reachable and responding while further analysis continued to isolate the source of the failure. \n3. Cause Identified: Investigation determined that a required service in the EU production environment had been unexpectedly stopped. \n4. Service Restarted: The affected service was restarted in the EU production environment. \n5. Upload Recovery Validated: Technical teams confirmed that uploads were processing successfully again following service restoration. \n6. Post-Restoration Monitoring Performed: Recovery activity and backlog behavior were monitored after restoration.   \n\n**Future Preventative Measures**\n\nThe following follow-up actions were identified during the incident:\n\n1. Automatic Service Restart Measures: We are implementing measures to help ensure the affected service starts automatically in the event of a similar occurrence.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-15T12:28:18Z",
"resolved_at": "2026-04-08T07:11:01.010-07:00",
"resolved_inferred": false,
"started_at": "2026-04-08T04:07:21.487-07:00",
"state": "postmortem",
"title": "Flexera One \u2013 IT Asset Management \u2013 EU \u2013 Inventory Uploads Failures",
"updated_at": "2026-04-21T20:44:25.909-07:00",
"url": "https://stspg.io/r880cz91xm1x"
},
{
"body": "**Description:** CloudCheckr \u2013 All Regions \u2013 Inventory and Cost Collection Delays\n\n**Timeframe:** April 1, 2026, 1:00 PM PDT \u2013 April 6, 2026, 8:00 PM PDT\n\n**Incident Summary**\n\nOn April 1, 2026, at approximately 1:00 PM PDT, an issue was identified affecting inventory and cost collection within CloudCheckr across all regions. During this period, some customers experienced delays in inventory and cost data collection.  \n  \nDuring the investigation, technical teams confirmed impact across US, EU, AU, GOV, and HSE regions. The issue was linked to service disruptions affecting an external cloud service provider in Middle East regions. As the incident progressed, technical teams confirmed that the issue was causing failures in discovery workflows and was also impacting billing and invoicing. For customers with usage in the affected regions, collection workflows were delayed while processing waited on timeouts before moving on.  \n  \nTechnical teams monitored the environment closely while developing and validating a mitigation. A product-side mitigation was then deployed across all CloudCheckr regions to bypass the affected Middle East region on a per-customer basis when communication with that region failed. Following completion of the deployments, technical teams continued monitoring and validating service behavior across regions.   \n  \nBy April 6, 2026, at approximately 8:00 PM PDT, recovery had been confirmed, with no remaining concerns at that time regarding data processing, billing, or invoicing, and the incident was considered resolved.\n\n**Root Cause**    \nThe incident was caused by service disruptions affecting an external cloud service provider in Middle East regions. As a result, CloudCheckr collection workflows encountered timeouts when processing customers with usage in the affected regions, which led to delays in inventory and cost data collection and also impacted billing and invoicing.\n\n**Remediation Actions**    \nThe following actions were taken during the incident response:\n\n1. Incident Investigation Initiated: Technical teams began investigating delays affecting inventory and cost collection across CloudCheckr regions.\n2. External Service Disruption Identified: The issue was linked to service disruptions affecting an external cloud service provider in Middle East regions.\n3. Impact Assessment Performed: Technical teams assessed the effect on collection workflows, including discovery, billing, and invoicing.\n4. Monitoring Continued: Teams continued to monitor service behavior and customer impact while evaluating mitigation options.\n5. Product-Side Mitigation Deployed: A mitigation was deployed across CloudCheckr regions to bypass the affected Middle East region on a per-customer basis when communication with that region failed.\n6. Recovery Validated: Following deployment, technical teams monitored the environment and confirmed recovery across all regions before resolving the incident.\n\n**Future Preventative Measures**\n\n\u2022\tResilient Regional Collection Handling: We will further strengthen CloudCheckr\u2019s collection logic to reduce dependency on a single affected region during external service disruptions. This includes improving how collection workflows detect repeated regional communication failures and continue processing in a way that helps minimize delays to inventory, cost, billing, and invoicing activities across other unaffected regions.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-14T12:29:13Z",
"resolved_at": "2026-04-07T14:36:03.334-07:00",
"resolved_inferred": false,
"started_at": "2026-04-02T13:50:56.000-07:00",
"state": "postmortem",
"title": "CloudCheckr \u2013 All Regions \u2013 Inventory and Cost Collection Delays",
"updated_at": "2026-04-30T21:47:48.001-07:00",
"url": "https://stspg.io/8lfvc9lhyq8g"
},
{
"body": "**Description:** Flexera One \u2013 NAM \u2013 Access Disruption\n\n**Timeframe:** April 2, 2026, 12:39 PM PDT \u2013 April 2, 2026, 1:02 PM PDT\n\n## Incident Summary\n\nOn April 2, 2026, at approximately 12:39 PM PDT, an issue was identified affecting access to the Flexera One application in the North America \\(NAM\\) region. During this period, customers may have experienced difficulties accessing the Flexera One application.\n\nTechnical teams began investigating immediately and identified elevated load affecting access-related backend services involved in authentication and request processing. This degraded service behavior temporarily impacted customer access in the NAM region.\n\nMitigation actions were initiated during the incident response, including reverting recent changes associated with the affected services. Following these actions, system performance improved and customer access was restored.\n\nBy April 2, 2026, at approximately 1:02 PM PDT, access to the Flexera One application in NAM had returned to expected operation. After validation and monitoring confirmed stable recovery, the incident was considered resolved.\n\n## Root Cause\n\nThe incident was caused by a compounding database performance issue involving an existing database request pattern and a recent service change affecting access-related request handling. Together, these conditions created elevated load in the production environment and temporarily disrupted customer access to the Flexera One application.\n\n## Remediation Actions\n\nThe following actions were taken during the incident response:\n\n1. **Incident Detection and Response Initiated:** Technical teams were alerted to the access disruption affecting the NAM region and began immediate investigation.\n2. **Impact Isolation:** It was confirmed that the customer-facing impact was limited to the NAM region.\n3. **Service Review and Diagnosis:** Technical teams reviewed recent changes and identified elevated load affecting access-related backend services.\n4. **Corrective Action Applied:** Recent changes associated with the affected services were reverted to restore normal operation.\n5. **Service Restoration Verification:** Customer access and application performance were validated following the mitigation actions.\n6. **Post-Recovery Monitoring:** Services were monitored after restoration to confirm continued normal operation.\n\n## Future Preventative Measures\n\nThis incident highlighted the importance of validating performance-related service changes and database-heavy access paths under higher production traffic conditions.\n\nBased on this experience, the following measures are being applied:\n\n1. **Performance Validation Enhancements:** We are reviewing pre-production validation approaches to better identify performance-related issues that may only surface under higher traffic conditions.\n2. **Load Testing Improvements:** We are enhancing load testing coverage to better validate database-heavy access paths using production-observed traffic patterns.\n3. **Database Access Optimization:** Technical teams have implemented query improvements and a fix related to the identified database performance behavior.\n4. **Service Health Check Review:** We are reviewing health check coverage for access-related services to improve visibility into customer-impacting access failures.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "critical",
"last_seen": "2026-09-10T12:15:40Z",
"resolved_at": "2026-04-02T13:47:12.519-07:00",
"resolved_inferred": false,
"started_at": "2026-04-02T13:02:30.862-07:00",
"state": "postmortem",
"title": "Flexera One - NAM - Access Disruption",
"updated_at": "2026-05-18T11:53:15.177-07:00",
"url": "https://stspg.io/kckxzgl252kp"
},
{
"body": "**Description:** Flexera One \u2013 IT Visibility \u2013 All Regions \u2013 Third-Party Inventory Import Processing Disruption\n\n**Timeframe:** March 31, 2026, 3:00 PM PDT to April 3, 2026, 11:12 PM PDT\n\n\u200c\n\n**Incident Summary**\n\n\u200c\n\nOn Wednesday, March 31, 2026, at 3:00 PM PDT, an issue was identified affecting third-party \\(external\\) inventory imports within Flexera One IT Visibility across all regions. During the impact window, customers using these external inventory connections experienced delays in inventory processing and timeouts during import operations. The issue was observed across all regions \\(NAM, EU, and APAC\\), while other inventory ingestion methods remained fully operational.\n\nMultiple customer cases were reported, and technical teams initiated an investigation promptly. The issue was traced to disruptions in the inventory processing pipeline, where uploaded data was not advancing through the workflow as expected.\n\nAfter the issue was identified, a hotfix was deployed, restoring normal data flow for newly incoming inventory. While new data began processing successfully, some customers continued to experience delays as previously missed processing events were analyzed and system backlogs were cleared.\n\nSubsequent validation confirmed that newer inventory data superseded earlier missed data, removing the need for full reprocessing. Services were fully validated and confirmed to be operating as expected prior to incident closure on 3 Apr 2026, at 11:12 PM PDT.\n\n\u200c\n\n**Root Cause**\n\n\u200c\n\nThe issue was caused by a missing configuration setting introduced during a recent release, which impacted inventory processing across all regions.\n\nThis configuration gap resulted in certain processing events not being triggered, preventing uploaded inventory data from progressing through the expected pipeline.\n\nContributing Factors:\n\n* Configuration Gap in Release: A required environment setting was not present in production following deployment.\n\n* Processing Pipeline Interruption: Missing triggers prevented inventory data from progressing to downstream systems.\n\n* Regional Impact Variations: While the issue affected all regions, some regions experienced additional symptoms such as delayed processing queues.\n\n* Detection Delay: The issue primarily affected asynchronous processing, which delayed immediate visibility.\n\n\u200c\n\n**Remediation Actions**\n\n\u200c\n\nThe following remediation steps were implemented to restore service functionality:\n\n* Deployed a hotfix to restore the missing configuration across all regions.\n\n* Restored normal processing of inventory data for all newly submitted uploads.\n\n* Investigated and validated the status of previously impacted data.\n\n* Confirmed that newer data submissions correctly updated downstream systems.\n\n* Monitored system performance to ensure stability and full recovery. \n\n\u200c\n\n**Future Preventative Measures**\n\n\u200c\n\n* Implement Enhanced Release Validation Controls: Ensure all required environment configurations are validated and present prior to production deployment.\n\n* Strengthen End-to-End Pipeline Monitoring: Introduce comprehensive monitoring across asynchronous processing flows to detect failures or delays in real time.\n\n* Enforce Configuration Consistency Across Regions: Establish safeguards to prevent configuration drift and ensure uniform deployments across all environments.\n\n* Processing Gap Detection Alerts: Review and upgrade alerting mechanisms to identify missing or delayed processing events within the ingestion pipeline.\n\n* Optimize Recovery and Backlog Handling Procedures: Improve recovery strategies to efficiently handle missed processing events and reduce impact from backlog accumulation.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-11T12:29:42Z",
"resolved_at": "2026-04-03T23:13:18.930-07:00",
"resolved_inferred": false,
"started_at": "2026-03-31T15:12:13.546-07:00",
"state": "postmortem",
"title": "Flexera One \u2013 IT Visibility \u2013 All Regions \u2013 Third-Party Inventory Import Processing Disruption",
"updated_at": "2026-04-19T22:29:05.363-07:00",
"url": "https://stspg.io/9wlp4nw8vx2y"
},
{
"body": "**Description:** Flexera One \u2013 IT Visibility \u2013 NA \u2013 Data Explorer Request Failures\n\n**Timeframe:** March 31, 2026, 12:41 PM PDT to March 31, 2026, 1:56 PM PDT\n\n\u200c\n\n**Incident Summary**\n\n\u200c\n\nOn Tuesday, March 31, 2026, at 12:41 PM PDT, our teams detected an issue affecting the Data Explorer feature within Flexera One IT Visibility in the NA region. Customers in this region encountered errors when using Data Explorer; requests failed to return results as expected and, in some cases, resulted in HTTP 500 errors, preventing successful query execution.\n\nThe impact was isolated to Data Explorer functionality in the NA region. Other regions, including EU and APAC, were not affected, and the broader Flexera One platform remained fully accessible throughout the incident.\n\nEngineering teams began investigating immediately and determined that the issue originated in a backend service responsible for processing Data Explorer queries. Initial analysis suggested a potential problem with request handling. Subsequent validation confirmed that customer requests were valid and that the failure was occurring service-side within this backend component.\n\nThe issue was resolved by reverting the affected configuration changes in the backend service, which restored service stability. Post-recovery validation confirmed that Data Explorer queries completed successfully and that normal functionality was fully restored for customers in the NA region.\n\n\u200c\n\n**Root Cause**\n\n\u200c\n\nDuring their investigations, our technical teams identified that the issue was caused by a partial or inconsistent deployment in the NA region. Configuration changes were applied without the corresponding service components being fully deployed. This mismatch caused runtime failures during query processing, resulting in HTTP 500 errors when Data Explorer queries were executed.The issue was not related to authentication or customer-submitted queries, despite initial error messages suggesting otherwise.\n\nContributing Factors:\n\n* Partial Deployment State: Configuration updates were applied without matching service binaries. \n\n* Service Runtime Failures: The mismatch led to failures during query generation in the backend service. \n\n* Misleading Error Messages: UI errors suggested request issues, which delayed precise identification of the service-side cause. \n\n* Regional Isolation: The issue was limited to NA due to differences in deployment state across regions.\n\n\u200c\n\n**Remediation Actions**\n\n\u200c\n\nThe following remediation steps were implemented to restore service functionality:\n\n* Reverted the affected configuration changes in the NA environment. \n\n* Restored alignment between deployed configuration and service components. \n\n* Validated successful execution of Data Explorer queries across affected organizations. \n\n* Monitored system performance to confirm stability post-recovery.\n\n\u200c\n\n**Future Preventative Measures**\n\n\u200c\n\n* Review Post-Deployment Consistency Checks: Ensure configuration and service components are aligned across all regions after deployment.\n\n* Enhance Error Handling and Messaging: Improve system behavior so backend failures are accurately reflected in user-facing error messages.\n\n* Strengthen Regional Deployment Monitoring: Expand monitoring coverage for critical services to detect issues earlier.\n\n* Improve Rollback Validation Processes: Establish more robust validation steps to ensure faster and more reliable recovery following rollback actions.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-08T12:29:58Z",
"resolved_at": "2026-03-31T14:38:37.402-07:00",
"resolved_inferred": false,
"started_at": "2026-03-31T13:08:15.000-07:00",
"state": "postmortem",
"title": "Flexera One \u2013 IT Visibility \u2013 NA \u2013 Data Explorer Request Failures",
"updated_at": "2026-04-19T23:07:57.082-07:00",
"url": "https://stspg.io/fq6dkzyzn5xm"
},
{
"body": "**Description:** Flexera One \u2013 IT Visibility \u2013 EU \u2013 Errors Affecting Evidence UI and Export Functions\n\n**Timeframe:** March 31, 2026, 7:48 AM PDT to March 31, 2026, 1:20 PM PDT\n\n**Incident Summary**\n\nOn Tuesday, March 31, 2026, 7:48 AM PDT, an issue was identified affecting IT Visibility \\(ITV\\) functionality in the EU region. Customers experienced errors when accessing the Evidence UI, as well as failures when performing ZIP exports and Query exports. Requests returned HTTP 503 errors, resulting in incomplete or unsuccessful operations. During the impact window, customers in the EU region experienced failures across key ITV workflows, including accessing the Evidence UI and performing ZIP and Query exports. While the Flexera One platform remained accessible, these specific functionalities did not operate as expected.\n\nThe issue was detected through monitoring alerts and internal investigation. Engineering teams from multiple groups engaged immediately and worked in parallel to identify the cause and restore service. Initial mitigation efforts, including rollback of recent changes, did not fully resolve the issue. As part of recovery, a new infrastructure cluster was provisioned and traffic was redirected to it. Following this action, services began recovering, and full functionality was restored after DNS propagation completed.\n\nSome customers may have experienced brief residual impact due to caching before full recovery was realized.\n\n\u200c\n\n**Root Cause**\n\nThe issue was caused by a deployment-related configuration inconsistency affecting secure communication settings in the EU region.\n\nDuring a recent infrastructure deployment, a critical configuration responsible for secure service communication was unintentionally recreated. This led to intermittent failures in request routing, resulting in HTTP 503 errors for affected ITV functionalities.\n\nAlthough the deployment initially appeared successful, the issue manifested under specific conditions and impacted multiple organizations within the EU region.\n\nContributing Factors:\n\n* Configuration difference: Differences between deployed configurations in regions led to inconsistent behavior, with EU being uniquely impacted. \n* Deployment Side Effects: Infrastructure changes unintentionally modified critical communication settings. \n* Limited Functional Monitoring: Existing monitoring focused on service health, delaying detection of user-facing issues. \n\n\u200c\n\n**Remediation Actions**\n\n\u200c\n\nThe following remediation steps were implemented to restore service functionality:\n\n* Rolled back the impacted deployment changes in the EU region.\n* Provisioned a new infrastructure cluster and redirected traffic to stabilize services.\n* Verified recovery of affected UI, ZIP exports, and Query export functionality.\n* Monitored system behavior and confirmed stability configuration update propagation.\n\n\u200c\n\n**Future Preventative Measures**\n\n\u200c\n\n* Stronger Deployment Consistency Controls: Ensuring configuration changes are applied consistently across all regions to prevent drift.\n* Enhanced Post-Deployment Validation: Review and implement  functional \\(end-to-end\\) validation tests to verify key workflows after deployments.\n* Improved Monitoring and Alerting: Expanding monitoring to detect user-impacting failures \\(such as API and export failures\\).\n* Deployment Safeguards and Review: Strengthening change review processes to identify and prevent unintended configuration changes during deployments.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-08T12:29:58Z",
"resolved_at": "2026-03-31T13:44:35.239-07:00",
"resolved_inferred": false,
"started_at": "2026-03-31T08:04:19.000-07:00",
"state": "postmortem",
"title": "Flexera One \u2013 IT Visibility \u2013 EU \u2013 Errors Affecting Evidence UI and Export Functions",
"updated_at": "2026-04-15T23:42:12.169-07:00",
"url": "https://stspg.io/dhyd199nv06x"
}
]
}