{
"vendor": "Sauce Labs",
"slug": "sauce-labs",
"platform": "statuspage",
"status_url": "https://status.saucelabs.com",
"last_checked": "2026-09-16T12:28:20Z",
"last_state": "ok",
"history_backfilled": true,
"first_watched": "2026-09-04T07:06:16Z",
"incidents": [
{
"body": "Between 09:52 UTC and 10:34 UTC, multiple Sauce Labs services in the US-West data center experienced an outage, causing authentication failures and disruptions to testing features. All systems are now fully operational.",
"first_seen": "2026-09-16T12:28:20Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-09-16T11:08:02.493Z",
"resolved_inferred": false,
"started_at": "2026-09-16T11:08:02.460Z",
"state": "resolved",
"title": "2026-September-16 Service Incident",
"updated_at": "2026-09-16T11:08:02.501Z",
"url": "https://stspg.io/11s0hws4ft91"
},
{
"body": "After taking remedial action, Real Device test error rates have returned to normal in the EU-Central-1 data center. All services are fully operational.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-09-02T11:53:14.350Z",
"resolved_inferred": false,
"started_at": "2026-09-02T09:39:15.089Z",
"state": "resolved",
"title": "2026-September-02 Service Incident",
"updated_at": "2026-09-02T11:53:14.367Z",
"url": "https://stspg.io/7gpvwzpnd6r5"
},
{
"body": "### **Dates:**\n\nTuesday August 25th 2026, 11:38 \u2013 13:10 UTC\n\n### **What happened:**\n\nCustomers running macOS and iOS tests in our US-West region experienced degraded service. Roughly 50% of the virtual Mac capacity in the region stopped accepting new tests, so tests either queued or failed to start. Remaining capacity came under additional pressure as work shifted onto it, which extended start times for both desktop and simulator tests.\n\n### **Why it happened:**\n\nAn internal security certificate used by our Mac hosts to reach a supporting cloud service reached its expiry date. Once it lapsed, the hosts could no longer establish a trusted connection to that service and stopped provisioning new test machines. The certificate had been issued manually and had neither automated renewal nor expiry alerting.\n\n### **How we fixed it:**\n\nWe issued and deployed a replacement certificate, which restored connectivity and returned Mac capacity to normal levels.\n\n### **What we are doing to prevent it from happening again:**\n\nWe are moving these certificates onto automated renewal and adding alerting so they are replaced well ahead of expiry, along with reviewing the surrounding tooling to make certificate handling safer.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-08-25T13:19:04.761Z",
"resolved_inferred": false,
"started_at": "2026-08-25T12:33:24.168Z",
"state": "postmortem",
"title": "2026-August-25 Service Incident",
"updated_at": "2026-09-03T20:24:55.492Z",
"url": "https://stspg.io/2t9ydqgksldm"
},
{
"body": "### **Dates:**\n\nFriday, August 7th 2026, 11:00 UTC - 13:04 UTC.\n\n### **What happened:**\n\nWindows and Intel Mac jobs in `us-west1` failed to start due to virtual machine \\(VM\\) allocation starvation.\n\n### **Why it happened:**\n\nA service crash loop left VMs in an allocated but unclaimed state, while a cleanup bug prevented the system from releasing the orphaned capacity to boot new VMs.\n\n### **How we fixed it:**\n\nRestored service stability and cleared stale allocations to resume VM provisioning and clear queued jobs.\n\n### **What we are doing to prevent it from happening again:**\n\nFixing allocator cleanup logic, strengthening deployment health checks, and improving capacity accounting for stale allocations.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-08-07T13:06:06.442Z",
"resolved_inferred": false,
"started_at": "2026-08-07T12:08:31.103Z",
"state": "postmortem",
"title": "2026-August-07 Service Incident",
"updated_at": "2026-08-14T16:45:13.965Z",
"url": "https://stspg.io/sq0702yxhnqg"
},
{
"body": "### **Dates:**\n\nThursday July 16th 2026, 13:03 \u2013 17:33 UTC\n\n### **What happened:**\n\nSymbol archives uploaded in multiple parts failed to process. All regions were affected.\n\n### **Why it happened:**\n\nOur symbol processing service sent an upload-verification field that our cloud storage provider's API does not accept for multi-part uploads, so those uploads were rejected. A fix for this had already been developed, but it had not yet been included in a released build and the service was not configured to use it. As in the first incident, rejected uploads were retried and accumulated on local disk.\n\n### **How we fixed it:**\n\nWe deployed a build containing the fix and corrected the service configuration on the affected workers. A large backlog of uploads then processed, which briefly re-filled the disks before draining completely.\n\n### **What we are doing to prevent it from happening again:**\n\nWe are releasing the fix formally through our build pipeline and persisting the corrected configuration in our configuration management, so it can't be lost. The disk-utilization alerting added after the first incident also covers the disk-exhaustion pattern common to both.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-16T13:00:00.000Z",
"resolved_inferred": false,
"started_at": "2026-07-16T13:00:00.000Z",
"state": "postmortem",
"title": "2026-July-16 Resolved Service Incident - Error Reporting (Backtrace)",
"updated_at": "2026-08-14T16:49:35.724Z",
"url": "https://stspg.io/vs92t5j754q8"
},
{
"body": "### **Dates:**\n\nTuesday July 14th 2026, 23:45 UTC \u2013 Wednesday July 15th 2026, 23:05 UTC\n\n### **What happened:**\n\nSymbol archive uploads to Error Reporting \\(Backtrace\\) projects failed with HTTP 400 errors. All regions were affected.\n\n### **Why it happened:**\n\nThe credentials our symbol processing service used to write to cloud storage were no longer valid, so uploads could not be stored. Failed uploads were retried repeatedly and accumulated on local disk until the service ran out of space, at which point it also began rejecting new uploads.\n\n### **How we fixed it:**\n\nWe reissued the storage credentials, increased disk capacity on the affected workers, and restarted the service. The queued uploads then processed successfully, and we confirmed recovery with affected customers.\n\n### **What we are doing to prevent it from happening again:**\n\nWe've added disk-utilization monitoring and alerting to this service so we detect the condition ourselves before it affects uploads, rather than relying on customer reports.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-14T23:30:00.000Z",
"resolved_inferred": false,
"started_at": "2026-07-14T23:30:00.000Z",
"state": "postmortem",
"title": "2026-July-15 Resolved Service Incident - Error Reporting (Backtrace)",
"updated_at": "2026-08-14T16:48:21.879Z",
"url": "https://stspg.io/lbr0cfj9sp0s"
},
{
"body": "### **Dates:**\n\nTuesday July 14th 2026, 09:00 - 18:28 UTC\n\n### **What happened:**\n\niOS ARM tests running in our EU data center failed.\n\n### **Why it happened:**\n\nA fault occurred with our primary network provider in our EU data center.\n\n### **How we fixed it:**\n\nWe failed over to a secondary network provider from our EU data center.\n\n### **What we are doing to prevent it from happening again:**\n\nWe've improved our monitoring and alerting to catch issues with third party network providers.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-14T20:23:34.625Z",
"resolved_inferred": false,
"started_at": "2026-07-14T17:19:52.472Z",
"state": "postmortem",
"title": "2026-July-14 Service Incident",
"updated_at": "2026-07-29T17:46:11.189Z",
"url": "https://stspg.io/vk8vqbfz6bqb"
},
{
"body": "### **Dates:**\n\nWednesday July 1st 2026, 17:02 UTC - 21:31 UTC\n\n### **What happened:**\n\nmacOS 14 tests in the US West and EU Central data centers were unable to start.\n\n### **Why it happened:**\n\nAn internal datasource was unavailable due to a missing configuration entry.\n\n### **How we fixed it:**\n\nThe missing entry was replaced.\n\n### **What we are doing to prevent it from happening again:**\n\nSafeguards have been put in place to prevent in-use datasource removal from configuration.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-01T17:02:00.000Z",
"resolved_inferred": false,
"started_at": "2026-07-01T17:02:00.000Z",
"state": "postmortem",
"title": "2026-July-1 Resolved Service Incident 1",
"updated_at": "2026-08-11T20:46:45.841Z",
"url": "https://stspg.io/xtz3j0b48605"
},
{
"body": "### **Dates:**\n\nWednesday July 1st 2026, 09:37 \u2013 12:46 UTC\n\n### **What happened:**\n\nThe Appium Inspector feature, used during real device live testing sessions, became unavailable after a scheduled UI deployment. Users who attempted to use the feature during a live test were presented with a 500 error, disrupting their active testing session. Automated test pipelines were not affected.\n\n### **Why it happened:**\n\nA routine upgrade of a core frontend library introduced an incompatibility with the Appium Inspector's rendering logic. The previous library version tolerated the pattern, but the updated version did not. The issue was not caught before deployment due to insufficient end-to-end test coverage for this specific feature.\n\n### **How we fixed it:**\n\nWe rolled back the UI to the last known working version to restore the feature immediately, and then deployed a targeted fix for the incompatibility.\n\n### **What we are doing to prevent it from happening again:**\n\nWe are adding end-to-end test coverage for the Appium Inspector feature to ensure it is validated automatically before future deployments.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-01T12:58:16.164Z",
"resolved_inferred": false,
"started_at": "2026-07-01T12:12:33.438Z",
"state": "postmortem",
"title": "2026-July-01 Service Incident",
"updated_at": "2026-07-20T21:01:42.383Z",
"url": "https://stspg.io/rgrqs3y0dwnx"
},
{
"body": "### **Dates:** \n\nWednesday, June 24th 2026, 09:42 UTC - 15:45 UTC.\n\n### **What happened:**\n\nReal Device test sessions using Appium and Access API experienced increased error rates in the US East, US West, and EU data centers.\n\nSessions were timing out after approximately 90 seconds or becoming stuck in the \"Connecting\" state, preventing them from being closed.\n\n### **Why it happened:**\n\nA product defect was introduced, causing Real Device test sessions to fail to start and preventing active sessions from being closed.\n\n### **How we fixed it:**\n\nRollback to a stable version.\n\n### **What we are doing to prevent it from happening again:**\n\nImprove monitoring & alerting and enhance post deployment validation.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-06-24T16:39:09.889Z",
"resolved_inferred": false,
"started_at": "2026-06-24T16:39:09.847Z",
"state": "postmortem",
"title": "2026-June-24 Resolved Service Incident",
"updated_at": "2026-07-22T20:50:58.378Z",
"url": "https://stspg.io/fvy9cvc0wq3s"
},
{
"body": "### **Dates:**\n\nThursday June 18 2026, 03:45 UTC - 06:45 UTC.\n\n### **What happened:**\n\nApproximately 7% of iOS devices in our EU data center were temporarily unavailable for customer test sessions after a loss of power to the rack hosting them.\n\n### **Why it happened:**\n\nThe power circuit feeding the affected rack exceeded its capacity, and a protective breaker tripped to safeguard the line - cutting power to the devices on that rack until the circuit was restored.\n\n### **How we fixed it:**\n\nThe affected circuit was reset and power restored to the rack, returning the devices to customer service.\n\n### **What we are doing to prevent it from happening again:**\n\nWe are redistributing power load across affected racks, adding capacity monitoring with early-warning alerts ahead of circuit limits, and introducing a capacity review before new devices are deployed to a rack.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-06-18T06:59:38.145Z",
"resolved_inferred": false,
"started_at": "2026-06-18T05:30:39.270Z",
"state": "postmortem",
"title": "2026-June-18 Service Incident",
"updated_at": "2026-07-06T17:32:59.147Z",
"url": "https://stspg.io/bqy2th03pgdz"
},
{
"body": "### **Dates:**\n\nMonday, June 1st 2026, 13:06 UTC -\u00a0 15:42 UTC.\n\n### **What happened:**\n\nCustomers served by the US-EAST-4 region were unable to authenticate or start new test sessions because an incomplete TLS certificate chain was deployed to the core directory services.\u00a0\n\n### **Why it happened:**\n\nA certificate extraction script defect silently truncated the certificate chain after the certificate authority transitioned to a longer hierarchy.\u00a0\n\n### **How we fixed it:**\n\nReverted the certificate rotation and re-applied the previous known ,good certificate to restore authentication services.\n\n### **What we are doing to prevent it from happening again:**\n\nImplementing pre-deployment certificate chain validation, adding active monitoring, and fixing the script's chain-length limitations.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "critical",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-06-01T18:24:32.108Z",
"resolved_inferred": false,
"started_at": "2026-06-01T14:02:09.100Z",
"state": "postmortem",
"title": "2026-June-1 Service Incident",
"updated_at": "2026-06-15T23:32:55.716Z",
"url": "https://stspg.io/7l70v858zpwm"
},
{
"body": "### **Dates:**\n\nWednesday, May 13th 2026, 18:47 UTC - Friday, May 15th 2026, 09:23 UTC.\n\n### **What happened:**\n\nCustomers were unable to access test run summaries for Real Device Cloud \\(RDC\\) jobs because events stopped publishing to the jobs Kafka topic in US-EAST.\n\n### **Why it happened:**\n\nAn authentication key used by the message producer unexpectedly lost its permissions during an account cleanup.\n\n### **How we fixed it:**\n\nManually restored the required permissions to re-establish the connection and resume service.\n\n### **What we are doing to prevent it from happening again:**\n\nMigrating to a permanent service account, implementing a Dead Letter Queue \\(DLQ\\) for the jobs Kafka topic, and replaying the missing events to restore customer data.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-05-21T14:53:43.957Z",
"resolved_inferred": false,
"started_at": "2026-05-21T14:53:43.890Z",
"state": "postmortem",
"title": "2026-May-13 Resolved Service Incident",
"updated_at": "2026-05-26T15:08:03.794Z",
"url": "https://stspg.io/55lgnt55wjd4"
},
{
"body": "### **Dates:**\n\nThursday, April 23rd 2026, 22:43 UTC - Friday, April 24th 2026, 15:29 UTC\n\n### **What happened:**\n\nVideo assets were missing for virtual iOS simulator tests on ARM and macOS ARM desktop tests in the US-West and EU data centers.\n\n### **Why it happened:**\n\nA product defect was introduced resulting in a screen capture failure.\n\n### **How we fixed it:**\n\nWe performed a rollback to a stable version.\n\n### **What we are doing to prevent it from happening again:**\n\nWe are improving monitoring & alerting to enhance our post deployment validation.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-04-24T16:33:12.764Z",
"resolved_inferred": false,
"started_at": "2026-04-24T16:33:12.727Z",
"state": "postmortem",
"title": "2026-April-23 Resolved Service Incident",
"updated_at": "2026-05-13T15:10:36.577Z",
"url": "https://stspg.io/qzrzvgdd9srp"
},
{
"body": "### **Dates:**\n\nThursday, April 16th 2026, 00:00 UTC \u2013 09:15 UTC\n\n### **What happened:**\n\nLive and automated tests on iOS 17.0 simulators failed to start in both the EU and US-West data centers. Customers running tests on iOS 17.0 Intel-based simulators were unable to execute their tests for approximately 9 hours.\n\n### **Why it happened:**\n\nA deployment introduced an incompatibility affecting iOS 17.0 on Intel-based infrastructure. The issue was not caught prior to release due to insufficient post-deployment test coverage for that specific simulator configuration.\n\n### **How we fixed it:**\n\nWe performed a rollback to the previous deployment, which restored full iOS 17.0 simulator functionality.\n\n### **What we are doing to prevent it from happening again:**\n\nWe are reviving and expanding automated post-deployment tests to cover a broader range of simulator configurations, including legacy Intel-based iOS versions, to catch incompatibilities before they reach production.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-04-16T10:10:56.208Z",
"resolved_inferred": false,
"started_at": "2026-04-16T10:10:56.165Z",
"state": "postmortem",
"title": "2026-April-16 Resolved Service Incident",
"updated_at": "2026-05-08T18:59:45.139Z",
"url": "https://stspg.io/kd3z9szx8n3c"
},
{
"body": "### **Dates:**\n\nMonday April 7th 2026, ~11:00 \u2013 15:55 UTC\n\n### **What happened:**\n\nSome customers experienced 503 errors when running tests via saucectl. The test-composer service was intermittently unavailable, preventing framework-based test execution.\n\n### **Why it happened:**\n\nA stale Docker image was deployed to the test-composer service due to a packaging issue that arose during an internal container registry migration. This caused service pods to crash.\n\n### **How we fixed it:**\n\nWe identified the stale image and redeployed the correct version, restoring the service.\n\n### **What we are doing to prevent it from happening again:**\n\nWe are hardening our image deployment pipeline and adding validation checks to ensure container registry migrations do not result in stale or incorrect images being deployed to production.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-04-07T16:13:09.918Z",
"resolved_inferred": false,
"started_at": "2026-04-07T15:02:39.190Z",
"state": "postmortem",
"title": "2026-April-07 Service Incident",
"updated_at": "2026-04-10T21:58:06.588Z",
"url": "https://stspg.io/x33t24l825cq"
},
{
"body": "### **Dates:**\n\nTuesday, March 24th 2026, 09:32 UTC \u2013 15:13 UTC\n\n### **What happened:**\n\nNetwork calls failed on iOS devices during Real Device Cloud sessions where network capture was enabled. Approximately 12-13% of iOS sessions were affected. Android was not impacted.\n\n### **Why it happened:**\n\nA deployment introduced a DNS resolution change that was incompatible with the iOS platform, causing network capture to break.\n\n### **How we fixed it:**\n\nRolled back the deployment to restore service.\n\n### **What we are doing to prevent it from happening again:**\n\nAdding synthetic tests to catch network capture regressions before production, and implementing monitoring alerts for faster detection after deployments.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-03-24T17:36:32.823Z",
"resolved_inferred": false,
"started_at": "2026-03-24T17:36:32.770Z",
"state": "postmortem",
"title": "2026-March-24 Resolved Service Incident",
"updated_at": "2026-04-10T23:07:27.712Z",
"url": "https://stspg.io/7t8yy3cvkzsb"
},
{
"body": "### **Dates:**\n\nWednesday, March 19 2026, 04:45 UTC - 10:47 UTC.\n\n### **What happened:**\n\nApproximately 15% of iOS devices in our US-West data center were temporarily unavailable for customer test sessions due to failed internet connectivity checks.\n\n### **Why it happened:**\n\nAn automated wireless network optimization feature adjusted transmit power levels on access points serving the affected devices, degrading wireless connectivity and causing devices to fail their availability checks.\n\n### **How we fixed it:**\n\nThe affected access points were identified and restarted, restoring normal wireless connectivity.\n\n### **What we are doing to prevent it from happening again:**\n\nEvaluation of the automated optimization tools and a monitoring improvement.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-03-19T10:54:16.140Z",
"resolved_inferred": false,
"started_at": "2026-03-19T09:51:49.053Z",
"state": "postmortem",
"title": "2026-March-19 Service Incident",
"updated_at": "2026-04-15T19:16:52.046Z",
"url": "https://stspg.io/qsrmlzlqy09n"
},
{
"body": "### **Dates:**\n\nFriday, March 13th 2026, 14:43 UTC - 15:11 UTC.\n\n### **What happened:**\n\nReal Devices \\(iOS and Android\\) availability gradually decreased across all data centers.\n\n### **Why it happened:**\n\nA product defect was introduced resulting in a small subset of Real Devices \\(~10%\\) failing to maintain required connectivity.\n\n### **How we fixed it:**\n\nRollback to a stable version.\n\n### **What we are doing to prevent it from happening again:**\n\nImprove monitoring & alerting, enhance post deployment validation.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-03-13T14:30:00.000Z",
"resolved_inferred": false,
"started_at": "2026-03-13T14:30:00.000Z",
"state": "postmortem",
"title": "2026-March-13 Resolved Service Incident",
"updated_at": "2026-04-10T22:58:23.241Z",
"url": "https://stspg.io/xl8wdzgbv98p"
},
{
"body": "### **Dates:**\n\nTuesday March 10th 2026, 17:52 - 23:34 UTC\n\n### **What happened:**\n\nThe majority of iOS devices across all regions became unavailable.\n\n### **Why it happened:**\n\nApple's [ppq.apple.com](http://ppq.apple.com) app verification endpoint was down, causing internal device monitoring checks to fail, bringing devices offline.\n\n### **How we fixed it:**\n\nWe temporarily disabled these device monitoring checks.\n\n### **What we are doing to prevent it from happening again:**\n\nImproved external monitoring to catch outages of apple\u2019s [ppq.apple.com](http://ppq.apple.com) endpoint, loosened device monitoring to not take down live iOS devices if [ppq.apple.com](http://ppq.apple.com) is down.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-03-10T23:34:47.448Z",
"resolved_inferred": false,
"started_at": "2026-03-10T18:46:06.000Z",
"state": "postmortem",
"title": "2026-March-10 Service Incident",
"updated_at": "2026-04-10T18:54:29.970Z",
"url": "https://stspg.io/2629b66924m5"
},
{
"body": "### **Dates:** \n\nFriday, March 6th 2026, 21:38 UTC - 23:11 UTC\n\n### **What happened:**\n\nDuring the incident timeline, customers running virtual iOS simulator tests on ARM or macOS ARM desktop tests in the EU Data Center were unable to start new sessions for either live or automated.\n\n### **Why it happened:**\n\nThere was a sequencing issue on the release of the ARM side disk images in the EU.\n\n### **How we fixed it:**\n\nThe image reference for the ARM side disk was rolled back to the previous reference to restore service.\n\n### **What we are doing to prevent it from happening again:**\n\nThe tests that run to validate the image syncing have been completed in each region.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-03-06T21:38:00.000Z",
"resolved_inferred": false,
"started_at": "2026-03-06T21:38:00.000Z",
"state": "postmortem",
"title": "2026-March-6 Resolved Service Incident",
"updated_at": "2026-04-10T22:55:57.296Z",
"url": "https://stspg.io/n3lfqk75kx30"
},
{
"body": "### **Dates:**\n\nFriday, February 27th 2026, 16:15 UTC - 19:50 UTC\u00a0\n\n### **What happened:**\n\nRequests made using API client authentication would return 500 errors.\u00a0\n\n### **Why it happened:**\n\nAn internal data structure became corrupted due to a race condition.\n\n### **How we fixed it:**\n\nThe affected service was restarted, and a long-term fix was applied.\u00a0\n\n### **What we are doing to prevent it from happening again:**\n\nThread locking has been applied to the affected service.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-02-27T20:15:31.802Z",
"resolved_inferred": false,
"started_at": "2026-02-27T18:42:01.000Z",
"state": "postmortem",
"title": "2026-February-27 Service Incident",
"updated_at": "2026-04-15T19:10:24.220Z",
"url": "https://stspg.io/6kt2bq70kmvp"
},
{
"body": "### **Dates:**\n\nMonday, February 23rd 2026, 17:22 UTC - 22:31 UTC\n\n### **What happened:**\n\nWDIO-based tests run by customers using Sauce Connect 4 could not be started.\n\n### **Why it happened:**\n\nMisconfiguration caused failing health checks in some cases, causing customer tunnels to shut down.\n\n### **How we fixed it:**\n\nMisconfiguration was corrected.\n\n### **What we are doing to prevent it from happening again:**\n\nEvaluation of the underlying software stack.\u00a0\n\nCustomers who are still using SC4 that can migrate to SC5 should do so at their earliest convenience.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-02-23T22:31:50.952Z",
"resolved_inferred": false,
"started_at": "2026-02-23T17:22:02.217Z",
"state": "postmortem",
"title": "2026-February-23 Service Incident",
"updated_at": "2026-04-15T22:32:38.436Z",
"url": "https://stspg.io/v7cttsc5jc76"
},
{
"body": "### **Dates:**\n\nThursday, January 29th 2026, 19:30 UTC - 22:25 UTC.\n\n### **What happened:**\n\nApplication uploads and automated tests in the US-West-1 data center experienced intermittent errors between 5:43 pm \\(UTC\\) and 9:45 pm \\(UTC\\)\n\n### **Why it happened:**\n\nDue to a high number of concurrent uploads and a temporary change to the network topology schema, the App Storage Service experienced increased connection latency and delays in acquiring connections to the backend.\n\n### **How we fixed it:**\n\nApp Storage Service network topology schema was restored to its original state.\n\n### **What we are doing to prevent it from happening again:**\n\nWe are improving monitoring, alerting and validations for the App Storage Service backend connectivity.\n\nWe are working to implement processes to handle periods of increased load and proactively detect performance degradation to implement automated remediation and self-healing capabilities.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-01-29T17:30:00.000Z",
"resolved_inferred": false,
"started_at": "2026-01-29T17:30:00.000Z",
"state": "postmortem",
"title": "2026 - January - 29  Resolved Service Incident",
"updated_at": "2026-03-06T13:48:45.541Z",
"url": "https://stspg.io/2q85hbtzw2v5"
},
{
"body": "### **Dates:**\n\nThursday, January 29th 2026, 07:04 - 08:46 UTC.\n\n### **What happened:**\n\nBetween 07:04 UTC and 08:46 UTC on January 29, 2026, customers using our US-West-1 region experienced instability with Mac resources. Specifically:\n\n* Approximately 20% of iOS simulator jobs failed to start due to infrastructure errors.\n\n* 10% of Virtual MacOS/iOS jobs failed to upload assets to S3. As a result, logs, videos, and screenshots for these specific jobs were lost and remain unavailable for download.\n\n### **Why it happened:**\n\nScheduled maintenance on a primary network circuit coincided with high traffic volume, which led to congestion on the backup circuit.\n\n### **How we fixed it:**\n\nFull connectivity was restored when traffic was routed back to the primary circuit following the completion of maintenance.\n\n### **What we are doing to prevent it from happening again:**\n\nReview network redundancy strategies and consider upgrading backup circuit capacity to ensure it can support peak traffic loads.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-01-29T11:11:02.903Z",
"resolved_inferred": false,
"started_at": "2026-01-29T11:11:02.860Z",
"state": "postmortem",
"title": "2026-January-29 Resolved Service Incident",
"updated_at": "2026-02-03T14:01:51.202Z",
"url": "https://stspg.io/2mb9wb4g2d9d"
},
{
"body": "### **Dates:**\n\nTuesday, January 13th 2025, 9:10 UTC - 10:37 UTC.\n\n### **What happened:**\n\nThe App Storage Service experienced high response times in the US-West-1 datacenter.\n\n### **Why it happened:**\n\nDue to a high number of concurrent uploads, the App Storage Service experienced increased connection latency and delays in acquiring connections to the backend.\n\n### **How we fixed it:**\n\nService was restored by restarting the application.\n\n### **What we are doing to prevent it from happening again:**\n\nWe are improving monitoring, alerting and validations for the App Storage Service to backend connectivity. We are also evaluating how to handle periods of increased load and proactively detect performance degradation to implement automated remediation and self-healing capabilities.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-01-13T10:50:23.000Z",
"resolved_inferred": false,
"started_at": "2026-01-13T10:43:07.792Z",
"state": "postmortem",
"title": "2026-January-13 Service Incident",
"updated_at": "2026-02-03T13:58:45.644Z",
"url": "https://stspg.io/c5jtb2fzrnq7"
},
{
"body": "### **Dates:**\n\nMonday, January 5th 2026, 20:40 UTC - 21:04 UTC.\n\n### **What happened:**\n\nThere was a decrease in Real Device availability in the US-West-1 datacenter.\n\n### **Why it happened:**\n\nDuring a planned traffic flow migration, a network policy applied to the new traffic path was more restrictive than intended and this resulted in a subset of Real Devices \\(~40%\\) failing to maintain required connectivity.\n\nAlthough the change followed standard change management and peer review processes, the issue was not identified prior to activation.\n\n### **How we fixed it:**\n\nThe network policy was corrected and affected services recovered.\n\n### **What we are doing to prevent it from happening again:**\n\nWe are reviewing our network change processes to provide earlier detection of unintended behaviour during planned changes.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-01-05T22:24:24.843Z",
"resolved_inferred": false,
"started_at": "2026-01-05T21:15:54.374Z",
"state": "postmortem",
"title": "2026-January-5 Service Incident",
"updated_at": "2026-02-20T11:05:53.229Z",
"url": "https://stspg.io/83fr90kc5pmc"
},
{
"body": "### **Dates:**\n\nThursday, December 18th 2025, 5:32 UTC - 7:09 UTC.\n\n### **What happened:**\n\nThe App Storage Service experienced high response times in the US-West-1 datacenter.\n\n### **Why it happened:**\n\nDue to a high number of concurrent uploads, the App Storage Service experienced increased connection latency and delays in acquiring connections to the backend.\n\n### **How we fixed it:**\n\nThe service was restored by restarting the application.\n\n### **What we are doing to prevent it from happening again:**\n\nWe are implementing improved monitoring, alerting and validations for App Storage Service backend connectivity.\n\nWe will also look to enhance processes to handle periods of increased load and detect performance degradation proactively, in order to implement automated remediation and self-healing capabilities.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-12-18T05:30:00.000Z",
"resolved_inferred": false,
"started_at": "2025-12-18T05:30:00.000Z",
"state": "postmortem",
"title": "2025-December-18 Resolved Service Incident",
"updated_at": "2026-02-20T10:20:42.601Z",
"url": "https://stspg.io/kn6l21v1dqlk"
},
{
"body": "### **When it happened:**\n\nWednesday, December 17th 2025, 10:37 UTC - Thursday, December 18th 2025, 11:59 UTC\n\n### **What happened:**\n\nApplications were unable to be downloaded via Mobile App Distribution for a subset of customers.\n\n### **Why it happened:**\n\nA product change combined with an increased and sustained load resulted in higher resource usage than anticipated, exceeding established limits for certain clients.\n\n### **How we fixed it:**\n\nThe service was rolled back to a stable version.\n\n### **What we are doing to prevent it from happening again:**\n\nThere are multiple improvements planned.\u00a0\n\n1. Improved monitoring & alerting,\u00a0\n2. Enhanced post-deployment validation\n3. Capacity planning forecasting\n4. More advanced communication where possible ahead of major changes\n5. Adding canary to the roll out process",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-12-17T10:30:00.000Z",
"resolved_inferred": false,
"started_at": "2025-12-17T10:30:00.000Z",
"state": "postmortem",
"title": "2025-December-17 Resolved Service Incident",
"updated_at": "2026-01-05T16:09:35.393Z",
"url": "https://stspg.io/t1bqjv1793c4"
},
{
"body": "### **Dates:**\n\nWednesday, November 19th 2025, 09:00 UTC \u2013 Thursday, November 20th 2025, 21:00 UTC\n\n### **What happened:**\n\nFollowing a scheduled maintenance deployment, a subset of customers utilising Single Sign-On \\(SSO\\) were unable to log in to the platform. While the application remained up and running, authentication requests for these specific accounts were rejected, preventing access to the application.\n\n### **Why it happened:**\n\nA major backend infrastructure upgrade aimed at improving performance and security, inadvertently caused the incident. The new authentication system's default settings were incompatible with existing customer SSO configurations, leading to valid login attempts being incorrectly rejected.\n\n### **How we fixed it:**\n\nTo restore immediate access, we reverted the authentication processing logic to its previous state, ensuring full compatibility with all existing customer configurations.\n\n### **What we are doing to prevent it from happening again:**\n\nWe are adding enhanced monitoring for authentication success rates and updating pre-deployment testing to cover more legacy SSO configurations for seamless compatibility.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-11-19T22:17:25.647Z",
"resolved_inferred": false,
"started_at": "2025-11-19T18:54:55.602Z",
"state": "postmortem",
"title": "2025-November-19 Service Incident",
"updated_at": "2026-01-05T17:40:21.473Z",
"url": "https://stspg.io/zh4q2kw13smh"
},
{
"body": "### **Dates:**\n\nMonday November 4 2025, 04:35 UTC - 08:20 UTC.\n\n### **What happened:**\n\nThe web UI for Insights and Test Results was intermittently unresponsive.\n\n### **Why it happened:**\n\nA data store of components of the UI experienced unexpectedly high load, resulting in an unresponsive service.\n\n### **How we fixed it:**\n\nThe issue was resolved without human intervention.\n\n### **What we are doing to prevent it from happening again:**\n\nWe have improved support for the 3rd party vendor's logging and metrics tools.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-11-04T11:10:09.158Z",
"resolved_inferred": false,
"started_at": "2025-11-04T08:02:39.103Z",
"state": "postmortem",
"title": "2025-November-04 Service Incident",
"updated_at": "2026-01-06T14:50:23.109Z",
"url": "https://stspg.io/nsbjq1skbn89"
},
{
"body": "### **Dates:**\n\nMonday November 3 2025, 12:40 UTC - 14:25 UTC.\n\n### **What happened:**\n\nThe Sauce Labs dashboard for Insights and Test Results was unresponsive, causing issues showing test results and related information.\n\n### **Why it happened:**\n\nA data store for some components of the dashboard experienced unusually high load, resulting in an unresponsive service.\n\n### **How we fixed it:**\n\nThe issue was resolved by increasing the resources for the affected data store service.\n\n### **What we are doing to prevent it from happening again:**\n\nSupport for the 3rd party vendor's logging and metrics tools has been improved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-11-03T16:00:00.000Z",
"resolved_inferred": false,
"started_at": "2025-11-03T16:00:00.000Z",
"state": "postmortem",
"title": "2025-November-3 Resolved Service Incident",
"updated_at": "2025-12-11T17:30:21.410Z",
"url": "https://stspg.io/r3bwph15zgs4"
},
{
"body": "### **Dates:**\n\nMonday, October 27th 2025, 15:27 UTC - Wednesday, October 29th 2025, 11:34 UTC.\n\n### **What happened:**\n\nReal Device Live Testing availability gradually decreased in the US-West-1 data center.\n\n### **Why it happened:**\n\nA product defect was introduced which caused Live Testing sessions on Real Devices to fail to start.\n\n### **How we fixed it:**\n\nWe performed a rollback to the previous working release.\n\n### **What we are doing to prevent it from happening again:**\n\nWe have improved monitoring & alerting and enhanced our post-deployment validation processes.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "critical",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-10-29T10:47:31.946Z",
"resolved_inferred": false,
"started_at": "2025-10-29T09:27:27.439Z",
"state": "postmortem",
"title": "2025-October-29 Service Incident",
"updated_at": "2025-12-11T17:10:23.108Z",
"url": "https://stspg.io/b2y00kzgxb6b"
},
{
"body": "### **Dates:**\n\nTuesday, October 28th 2025, 14:03 UTC - 15:18 UTC \\(primary incident\\)  \nTuesday, October 28th 2025, 15:18 UTC - Thursday, October 30th 2025, 8:25 UTC \\(~1% of devices impacted window\\)\n\n### **What happened:**\n\nTest executions using the 'stable' Appium version gradually decreased in the US-West-1, US-East-4, and EU-Central-1 data centers.\n\n### **Why it happened:**\n\nA product defect was introduced which caused tests using the stable Appium version to fail immediately upon creation.\n\n### **How we fixed it:**\n\nWe preformed a rollback to a previous version on the majority of devices on Tuesday, October 28th 2025, 15:18 UTC, With a full service restoration across all devices completed on Thursday, October 30th 2025, 8:25 UTC.\n\n### **What we are doing to prevent it from happening again:**\n\nWe have improved monitoring & alerting and enhanced our post-deployment validation processes.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-10-28T16:13:00.000Z",
"resolved_inferred": false,
"started_at": "2025-10-28T14:07:00.000Z",
"state": "postmortem",
"title": "2025-October-28 Resolved Service Incident",
"updated_at": "2025-12-11T16:40:21.414Z",
"url": "https://stspg.io/hkvg72j4b5dk"
},
{
"body": "### **Dates:**\n\nTuesday, October 28th 2025, 11:40 UTC - 14:05 UTC.\n\n### **What happened:**\n\nWindows based tests in the EU data center were failing on startup which led to a decrease in availability of Windows-based Virtual Machines.\n\n### **Why it happened:**\n\nA product defect was introduced which caused an issue with starting Windows-based VMs.  \n\n### **How we fixed it:**\n\nThe deployment was rolled back to a stable version.\n\n### **What we are doing to prevent it from happening again:**\n\nWe have improved monitoring & alerting and enhanced our post-deployment validation.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-10-28T14:22:24.473Z",
"resolved_inferred": false,
"started_at": "2025-10-28T12:53:08.484Z",
"state": "postmortem",
"title": "2025-October-28 Service Incident",
"updated_at": "2025-11-06T16:06:16.170Z",
"url": "https://stspg.io/hkxwg4q3vy8g"
},
{
"body": "### **Dates:**\n\nThursday, October 9th 2025, 10:28 UTC - Tuesday, November 11th 2025, 11:34 UTC.\n\n### **What happened:**\n\nVirtual Android testing availability gradually decreased in the US-West-1 data center, this led to increased error rates in starting sessions, The error rates remained high between November 4th and November 11th.\n\n### **Why it happened:**\n\nAn increased and sustained load on our service led to reduced availability of Android virtual VMs, increasing wait times to start sessions, which in turn led to an elevated number of failed sessions.\n\n### **How we fixed it:**\n\nThis was resolved by increasing infrastructure capacity and optimizing resource utilization.\n\n### **What we are doing to prevent it from happening again:**\n\nWe have improved monitoring & alerting, enhanced post-deployment validation and improved capacity planning.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-10-28T11:30:00.000Z",
"resolved_inferred": false,
"started_at": "2025-10-28T11:30:00.000Z",
"state": "postmortem",
"title": "2025-October-28 Resolved Service Incident 1",
"updated_at": "2025-12-11T16:41:14.821Z",
"url": "https://stspg.io/r4vzdt8sgt36"
},
{
"body": "### **Dates:**\n\nWednesday, October 1st 2025, 10:34 UTC - Thursday, October 2nd 2025, 22:22 UTC\n\n### **What happened:**\n\nVirtual Android and Visual tests using the \u201cAndroid GoogleAPI Emulator\u201d were failing if they contained conditioned that required the exact app dimensions.\n\n### **Why it happened:**\n\nA new default emulator was released including a setting that did not account for customer defined virtual screen resolutions.\n\n### **How we fixed it:**\n\nThe issue was resolved by rolling back the deployment.\n\n### **What we are doing to prevent it from happening again:**\n\nA new release was scheduled with an update to accommodate tests with customer defined screen resolution.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-10-01T01:00:00.000Z",
"resolved_inferred": false,
"started_at": "2025-10-01T01:00:00.000Z",
"state": "postmortem",
"title": "2025-October-1 Resolved Service Incident",
"updated_at": "2025-12-11T16:30:21.330Z",
"url": "https://stspg.io/8h6br8g5y86x"
},
{
"body": "Between 06:30 and 08:35 UTC we were seeing errors when selecting and starting Live and Automated Android Emulator tests in US-West-1 & EU-Central-1 Data Centers. This has now been resolved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-09-19T05:30:00.000Z",
"resolved_inferred": false,
"started_at": "2025-09-19T05:30:00.000Z",
"state": "resolved",
"title": "2025-September-19 Resolved Service Incident",
"updated_at": "2025-09-19T08:54:31.597Z",
"url": "https://stspg.io/5gv3vpw1rsj2"
},
{
"body": "### **Dates:**\n\nWednesday September 5th 2025, 13:50 UTC - 16:31 UTC\n\n### **What happened:**\n\nThe App Storage Service in the EU Central 1 data center experienced elevated response times, leading to timeout errors with uploading applications. \n\n### **Why it happened:**\n\nThe App Storage Service experienced connection issues with the backend.\n\n### **How we fixed it:**\n\nThe connection issues were resolved by restarting the affected application.\n\n### **What we are doing to prevent it from happening again:**\n\nWe have improved observability signals for the App Storage Service and implemented regular scheduled restarts of the application.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-09-03T17:10:22.361Z",
"resolved_inferred": false,
"started_at": "2025-09-03T16:34:57.323Z",
"state": "postmortem",
"title": "2025-September-3 Service Incident",
"updated_at": "2025-12-11T16:20:21.466Z",
"url": "https://stspg.io/3mx2m9j2ftsc"
},
{
"body": "### **Dates:**\n\nFriday August 19th 2025, 16:20 UTC - 16:55 UTC.\n\n### **What happened:**\n\nDuring the times stated, automated Android real device tests were failing to start intermittently in the US-West-1 datacenter.\n\n### **Why it happened:**\n\nA defect was introduced by a deployment which impacted Appium tests.\n\n### **How we fixed it:**\n\nThe deployment was rolled back.\n\n### **What we are doing to prevent it from happening again:**\n\nWe have improved our internal test cases to detect the conditions in which the defect occurred.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-08-19T19:00:00.000Z",
"resolved_inferred": false,
"started_at": "2025-08-19T19:00:00.000Z",
"state": "postmortem",
"title": "2025-August-19 Resolved Service Incident",
"updated_at": "2025-09-10T14:50:29.806Z",
"url": "https://stspg.io/mjc9cv8h5hly"
}
]
}