{
"vendor": "Keeper",
"slug": "keeper",
"platform": "statuspage",
"status_url": "https://statuspage.keeper.io",
"last_checked": "2026-09-16T12:28:20Z",
"last_state": "ok",
"history_backfilled": true,
"first_watched": "2026-09-04T07:06:16Z",
"incidents": [
{
"body": "This incident has been resolved.",
"first_seen": "2026-09-10T12:15:40Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-09-09T12:43:37.379-07:00",
"resolved_inferred": false,
"started_at": "2026-09-09T09:32:18.183-07:00",
"state": "resolved",
"title": "SMS 2FA issues with AWS aggregator",
"updated_at": "2026-09-09T12:43:37.397-07:00",
"url": "https://stspg.io/sgtmc5mmlktq"
},
{
"body": "# Post-Incident Report\n\n**July 31, 2026**\n\n## Summary\n\nOn July 30, 2026 at 11:33 PM CT, the KeeperPAM connection routing service in the EU region started throwing errors. The Router/Gateway errors lasted approximately 12 hours and 45 minutes, with full service restored at approximately 12:18 PM CT on July 31. During this window, the main Keeper EU platform \\(vault access, authentication, and all other Keeper services\\) remained fully operational throughout.\n\n## What Happened\n\nOn July 22, a routine deployment to our connection routing service \\(\u201cKeeper Router\u201d\\) included an updated dependency that changed how the service retrieves its startup configuration. A misconfiguration in the endpoint caused EU region to look for a FIPS endpoints, which are not available. As a result, when service containers in the EU region restarted following the deployment, they were unable to retrieve their startup configuration and entered a failed state.\n\nTwo factors allowed this failure to go undetected for over a week. First, existing containers continued serving traffic while replacement containers silently failed to start, so there was no immediate customer impact from the July 22 deployment. QA also passed all production verification tests. Second, the configuration-load failure was logged at INFO level rather than ERROR or CRITICAL, so no PagerDuty alerts were generated.\n\nOn the night of July 30, the last healthy containers cycled out, the service's load balancer had zero healthy targets, and the endpoint began returning HTTP 503 errors. Automated health-check monitoring detected the outage within seconds and paged our on-call team.\n\n## What We're Changing\n\n1. **Alerting on failed container rollovers.** We are deploying CloudWatch alarms across all regions and environments that fire when container tasks repeatedly fail to start, to detect silent deployment failures.\n2. **Log severity for startup failures.** Any failure to load required startup configuration will now log at ERROR/CRITICAL severity instead of INFO which trigger the proper actionable alerts.\n\nWe apologize to our EU customers for the disruption.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-30T23:00:00.000-07:00",
"resolved_inferred": false,
"started_at": "2026-07-30T23:00:00.000-07:00",
"state": "postmortem",
"title": "EU KeeperPAM router and gateway connections",
"updated_at": "2026-07-31T13:42:23.234-07:00",
"url": "https://stspg.io/sfd3sgm5mpbh"
},
{
"body": "API errors have been resolved. The Keeper Gateway and KSM client version is no longer returning an error.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-08T12:14:33.304-07:00",
"resolved_inferred": false,
"started_at": "2026-07-08T11:51:48.000-07:00",
"state": "resolved",
"title": "Resolved: Keeper Gateway and KSM API errors",
"updated_at": "2026-07-08T12:14:41.599-07:00",
"url": "https://stspg.io/4z131lx31smt"
},
{
"body": "The issue is resolved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "maintenance",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-06-26T16:26:11.756-07:00",
"resolved_inferred": false,
"started_at": "2026-06-26T15:43:05.888-07:00",
"state": "resolved",
"title": "GovCloud region maintenance",
"updated_at": "2026-06-26T16:26:11.773-07:00",
"url": "https://stspg.io/00z3qcqm8j47"
},
{
"body": "The issue was resolved at 12:30PM PST.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-06-19T16:00:38.540-07:00",
"resolved_inferred": false,
"started_at": "2026-06-19T08:05:50.000-07:00",
"state": "resolved",
"title": "Direct record sharing enforcement issue",
"updated_at": "2026-06-19T16:00:38.554-07:00",
"url": "https://stspg.io/9fn1hl2hblcd"
},
{
"body": "On March 30 at 8:00 PM PST, Keeper performed a scheduled minor upgrade to our multi-region AWS RDS clusters as part of routine maintenance. The upgrade itself completed in approximately 3 minutes, followed by a standard application restart process that took about 15 minutes. While this process is typically non-disruptive, recent changes to service health checks caused unexpected customer impact during the restart window.\n\nThis maintenance was not communicated in advance due to an internal miscommunication and the expectation of no impact. We apologize for the disruption and are taking steps to ensure all future maintenance\u2014regardless of expected impact\u2014is communicated ahead of time, as well as reviewing our health check configurations to prevent similar issues.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-03-30T20:00:00.000-07:00",
"resolved_inferred": false,
"started_at": "2026-03-30T20:00:00.000-07:00",
"state": "resolved",
"title": "API errors during maintenance operation",
"updated_at": "2026-03-31T18:41:16.080-07:00",
"url": "https://stspg.io/gv00jgz9gls4"
},
{
"body": "At 9:05 AM PST, alerts were triggered for issues affecting KeeperPAM connections managed through Keeper\u2019s ECS deployments in the US-EAST region. There were no recent changes to the application or environment.\n\nInvestigation identified a low-level concurrency bug in the Keeper EPM service that caused request failures under high simultaneous load. These failures led to instability in the ECS services supporting KeeperPAM connections.\n\nAs a temporary mitigation, we blocked the error condition, restoring KeeperPAM connectivity by approximately 1:00 PM PST.\n\nThe engineering team then developed and deployed an updated Keeper Router version to address the underlying issue and prevent EPM agents from triggering server errors. The fix was fully validated by 3:00 PM PST, at which point all KeeperPAM services were stable and operating normally.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-03-23T14:45:09.000-07:00",
"resolved_inferred": false,
"started_at": "2026-03-23T10:07:28.854-07:00",
"state": "postmortem",
"title": "Resolved: KeeperPAM Connection Errors in US Data Center",
"updated_at": "2026-03-23T16:25:24.451-07:00",
"url": "https://stspg.io/7pcmbl7mcjs5"
},
{
"body": "At 11:08 AM PST, DevOps and customers reported sporadic error messages occurring after login for KeeperPAM customers with connections and tunnels enabled. While vault login itself was not affected, the errors were triggered asynchronously after login due to a communication issue with the KeeperPAM router endpoint.\n\nFurther investigation determined that a \"scheduler\" database within the AWS RDS environment - used for certain PAM scheduling operations - was encountering an unexpected error. The operations team resolved the underlying database issue and restarted the affected services.\n\nThis scheduler service was inadvertently impacting connection establishment. To prevent similar issues in the future, we will be implementing software changes to ensure that scheduling-related errors cannot affect connection functionality.\n\nAdditional monitoring and alerting have also been implemented for the scheduler databases to help detect and prevent similar issues going forward.\n\nFull service for connection and tunneling capabilities was restored by 11:40 AM PST.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-03-16T11:52:01.031-07:00",
"resolved_inferred": false,
"started_at": "2026-03-16T11:25:39.762-07:00",
"state": "postmortem",
"title": "Resolved: KeeperPAM Connection Errors in US Data Center",
"updated_at": "2026-03-16T12:44:38.180-07:00",
"url": "https://stspg.io/xmv5k609bq2r"
},
{
"body": "Maintenance in the GovCloud region was scheduled to begin at 8:00 PM PST with a planned duration of 30 minutes.\n\nDuring the maintenance window, certain services required additional reconfiguration beyond the original scope. Core services were restored within the planned 30-minute window. However, due to the extended reconfiguration activities, some users experienced API errors during login for approximately 15 minutes following service restoration.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-02-18T21:30:15.729-08:00",
"resolved_inferred": false,
"started_at": "2026-02-18T21:30:15.665-08:00",
"state": "postmortem",
"title": "US GovCloud Region - Login Errors Resolved",
"updated_at": "2026-02-18T21:41:43.345-08:00",
"url": "https://stspg.io/ppmr21yzdrkv"
},
{
"body": "On December 2, our **EU region** experienced an unexpected outage during planned infrastructure updates. The issue began at **1:24 PM PT** and service was fully restored by **4:10 PM PT**.\n\n**What Happened**  \nDuring an update to our server configuration, a region-specific configuration error caused the EU environment to begin throwing API errors. Although the underlying issue was identified quickly, each corrective change required a full autoscaling group rollout \\(up to 30 minutes per cycle\\) which extended the total recovery time.\n\n**Resolution**  \nOur engineering team implemented the necessary fixes and restored service. We have also updated our infrastructure-as-code to ensure this issue cannot recur in future deployments.\n\nWe apologize for the disruption and appreciate your patience while we resolved the incident.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-12-02T16:47:56.077-08:00",
"resolved_inferred": false,
"started_at": "2025-12-02T14:48:46.000-08:00",
"state": "postmortem",
"title": "EU data center login errors - Resolved",
"updated_at": "2025-12-04T14:30:25.917-08:00",
"url": "https://stspg.io/s03qflw92x5v"
},
{
"body": "Clearing this alert as most AWS services are restored.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-10-20T13:31:17.449-07:00",
"resolved_inferred": false,
"started_at": "2025-10-20T10:25:40.746-07:00",
"state": "resolved",
"title": "AWS outage affecting KeeperPAM connections and scheduled rotations",
"updated_at": "2025-10-20T13:31:17.467-07:00",
"url": "https://stspg.io/09y8zn8d18l7"
},
{
"body": "All Keeper email delivery services in US-EAST-1 were restored around 4AM PST.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-10-20T08:17:46.979-07:00",
"resolved_inferred": false,
"started_at": "2025-10-20T02:20:49.722-07:00",
"state": "resolved",
"title": "Email delivery in US region due to AWS US-EAST outage",
"updated_at": "2025-10-20T08:17:46.995-07:00",
"url": "https://stspg.io/wylsd3hwpx38"
},
{
"body": "On **September 10 at 8:42 AM PDT**, our monitoring detected elevated errors with the **Auth API in the EU region**. The issue was identified promptly and fully resolved by **8:49 AM PDT**. Total impact duration was approximately **7 minutes**.\n\n### Impact\n\n* **Region affected:** EU\n* **Service affected:** Authentication API\n* **Customer impact:** Customers in the EU experienced failed authentication attempts during the incident window.\n\n### Timeline \\(PDT\\)\n\n* **08:42** \u2013 Incident detected. Elevated Auth API errors in the EU region reported.\n* **08:43** \u2013 Engineering and DevOps teams engaged.\n* **08:46** \u2013 Root cause isolated to a Terraform-related configuration change involving a VPC endpoint.\n* **08:49** \u2013 Configuration change rolled back; Auth API errors resolved.\n\n### Root Cause\n\nThe outage was caused by a configuration change in our production tenant during an ongoing Terraform migration. As part of the migration, a new VPC endpoint was being added to production environments. During Terraform apply, a subnet association was also being created. A network disruption occurred between the creation of the VPC endpoint and the subnet association \\(behavior not seen in prior testing\\) leading to temporary Auth API connectivity failures in the EU region. A new Terraform migration plan has been created that eliminates the possibility of disruption.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-09-10T08:49:23.691-07:00",
"resolved_inferred": false,
"started_at": "2025-09-10T08:42:41.658-07:00",
"state": "postmortem",
"title": "EU Region login API errors",
"updated_at": "2025-09-26T09:26:06.730-07:00",
"url": "https://stspg.io/p6vlghlxh5vy"
},
{
"body": "Resolved - At 9:49 AM PST, we experienced login API request failures due to errors on an RDS database in the US region. The database failover completed successfully, and all systems were fully operational by 10:00 AM PST.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-08-21T10:36:36.328-07:00",
"resolved_inferred": false,
"started_at": "2025-08-21T10:36:36.285-07:00",
"state": "resolved",
"title": "Resolved: Vault login API errors",
"updated_at": "2025-08-21T10:36:36.337-07:00",
"url": "https://stspg.io/3xx2lxq4wkvm"
}
]
}