{
"vendor": "Twingate",
"slug": "twingate",
"platform": "statuspage",
"status_url": "https://status.twingate.com",
"last_checked": "2026-09-16T12:28:20Z",
"last_state": "ok",
"history_backfilled": true,
"first_watched": "2026-09-04T07:06:16Z",
"incidents": [
{
"body": "This incident has been resolved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-09T00:24:46.552Z",
"resolved_inferred": false,
"started_at": "2026-07-08T22:04:03.678Z",
"state": "resolved",
"title": "Linux Package Downloads Down",
"updated_at": "2026-07-09T00:24:46.572Z",
"url": "https://stspg.io/702mqn92169c"
},
{
"body": "# Incident Report \u2013 Authorization Service Degradation\n\n## Components Impacted\n\nControl Plane \u2013 Authorization Service\n\n## Summary\n\nOn May 28, 2026, between 10:00 UTC and 12:30 UTC, Twingate experienced a degradation of its Authorization service that affected approximately 15% of active connections at peak impact.\n\nDuring the incident, authorization requests experienced elevated network latency. As request processing times increased, the Authorization service's effective capacity to handle incoming traffic was reduced, resulting in elevated error rates and intermittent authorization failures for a subset of customers.\n\nTwingate operates the Authorization service across multiple cloud regions in an active-active configuration. While the platform remained available throughout the event, the combination of increased request latency and reduced service capacity led to customer impact until additional capacity was provisioned and service performance stabilized.\n\n## Root Cause\n\nThe incident was triggered by elevated network latency affecting communication paths used by the Authorization service. As requests took longer to complete, individual service instances were able to process fewer requests than normal.\n\nThis reduction in throughput exposed a limitation in our auto-scaling configuration, which primarily relied on CPU utilization to determine service capacity requirements. As request-processing workers spent more time waiting on network operations, CPU utilization declined even as request latency increased. As a result, the service scaled down during a period of elevated request latency, reducing available capacity and amplifying customer impact.\n\nRecovery efforts were further complicated by an unusually high rate of spot instance preemptions in two regions, which reduced available compute capacity during stabilization.\n\n## Resolution\n\nEngineering teams mitigated the incident by manually increasing Authorization service capacity and expanding available cluster resources. As additional capacity came online, request latency and error rates returned to normal levels and service performance fully recovered.\n\n## Corrective Actions\n\n### Completed\n\n* Increased the baseline capacity of the Authorization service by raising the minimum number of service instances.\n* Increased baseline node capacity across affected Kubernetes clusters.\n* Implemented additional rate controls for unusually large authorization requests to better protect overall service availability during periods of elevated load.\n\n### In Progress\n\n* Rebalance spot and on-demand node capacity to reduce sensitivity to spot instance interruptions.\n* Enhance auto-scaling policies to incorporate latency and service performance metrics in addition to CPU utilization.\n* Expand monitoring and alerting to better detect conditions where request latency increases while resource utilization decreases.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "critical",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-05-28T17:27:10.677Z",
"resolved_inferred": false,
"started_at": "2026-05-28T09:24:48.603Z",
"state": "postmortem",
"title": "Twingate Service Incident",
"updated_at": "2026-06-03T16:41:30.297Z",
"url": "https://stspg.io/mlz24xnfv79y"
},
{
"body": "**Components impacted**\n\n* Data Plane - Relay 443 flows only\n\n**Summary**\n\nOn May 9, 2026, a subset of Twingate customers experienced an interruption to network access affecting Relay 443 connectivity. The disruption lasted approximately 2 hours and 10 minutes and was caused by a configuration error introduced during a routine software update. We have resolved the issue and are taking concrete steps to prevent a recurrence.\n\n**Root Cause**\n\nTwingate regularly releases software updates to our relay infrastructure, deploying them in stages across groups to minimize risk. As part of this update cycle, we also migrated our deployment tooling to a newer version of our infrastructure management system.\n\nThe newer tooling applies stricter configuration validation rules than its predecessor. While preparing the update, our team identified and corrected a configuration conflict this stricter validation had surfaced. However, the fix was shipped in the same release bundle as a second change that was designed to serve as a prerequisite \u2014 specifically, a change that prevents the stricter validation from modifying existing, live deployments. The result was that a critical network configuration \u2014 the definition that tells our relay nodes to accept connections on port 443 \u2014 was silently removed from existing cluster deployments.\n\n**Why wasn't the issue detected earlier with the initial rollout?**\n\nOur continuous smoke testing did not include end-to-end coverage for Relay 443 specifically, so no automated alert fired when previous groups were deployed. Because those early upgrade groups carried relatively low Relay 443 traffic, no customer-visible impact surfaced either, leaving the issue undetected for approximately 48 hours until the higher-traffic groups were rolled out.\n\n**Why wasn't the issue caught in lower environments?**\n\nThe issue was identified in a lower environment and a fix was prepared. However, the fix was deployed in the same release as the problematic change rather than as a prerequisite, which prevented it from taking full effect on existing cluster upgrades. Subsequent deployments confirmed the issue was fully resolved.\n\n**Remediation**\n\nOnce we confirmed the root cause, we re-deployed the corrected relay configuration \u2014 explicitly restoring the port 443 host mapping \u2014 across all cluster groups in sequence. We validated each group before proceeding to the next. Full recovery was confirmed at 11:55 UTC on May 9.\n\n**Corrective actions** Short-term:\n\n* **Completed:** Redeployed to production to confirm the fix is stable and the issue will not recur.\n* We are adding regional smoke tests for Relay 443 to match the coverage already in place for other relay deployments, ensuring failures like this are caught before impacting customers.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-05-08T10:55:24.442Z",
"resolved_inferred": false,
"started_at": "2026-05-08T08:44:42.777Z",
"state": "postmortem",
"title": "Relays443 Down",
"updated_at": "2026-05-14T09:21:02.498Z",
"url": "https://stspg.io/rq3prbfmff2g"
},
{
"body": "This incident has been resolved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-10-20T22:45:27.315Z",
"resolved_inferred": false,
"started_at": "2025-10-20T15:49:53.000Z",
"state": "resolved",
"title": "AWS Outage Impacting APT/RPM Package Repository (packages.twingate.com)",
"updated_at": "2025-10-20T22:45:27.332Z",
"url": "https://stspg.io/bhqwy6t6qzjz"
}
]
}