{
"vendor": "KnowBe4",
"slug": "knowbe4",
"platform": "statuspage",
"status_url": "https://status.knowbe4.com",
"last_checked": "2026-09-16T12:28:20Z",
"last_state": "ok",
"history_backfilled": true,
"first_watched": "2026-09-04T07:06:16Z",
"incidents": [
{
"body": "On Septermber 8th we made aware that inbound HTML emails rendered partially or blank in Classic Outlook due to a routine HTML library upgrade in Defend's harmful-code-removal engine interacting with Classic Outlook's legacy Word rendering engine. New Outlook and Outlook on the web were unaffected. This was strictly a display issue: no email was delayed, lost, or compromised, and threat detection remained fully active. Engineering has reverted the update to restore normal rendering. If you temporarily disabled harmful-code removal during this window, please re-enable it to ensure complete protection.",
"first_seen": "2026-09-12T12:30:15Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-09-08T07:00:00.000Z",
"resolved_inferred": false,
"started_at": "2026-09-08T07:00:00.000Z",
"state": "resolved",
"title": "Defend | Emails Not Being Rendered in Classic Outlook",
"updated_at": "2026-09-11T12:43:52.328Z",
"url": "https://stspg.io/p788lzy5p9z9"
},
{
"body": "This incident has been resolved.",
"first_seen": "2026-09-04T11:14:03Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-09-07T08:25:23.612Z",
"resolved_inferred": false,
"started_at": "2026-09-04T10:53:59.000Z",
"state": "resolved",
"title": "Workspace | Email MFA Verification Issues",
"updated_at": "2026-09-07T08:25:23.637Z",
"url": "https://stspg.io/qk3mw5plt3kg"
},
{
"body": "Our engineering team has implemented a fix for the issue that causing message processing delays in PhishER.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-08-26T19:51:45.044Z",
"resolved_inferred": false,
"started_at": "2026-08-26T16:19:52.231Z",
"state": "resolved",
"title": "PhishER - Message Processing Delays",
"updated_at": "2026-08-26T19:51:45.059Z",
"url": "https://stspg.io/lv4v7whd9w20"
},
{
"body": "On Friday, August 21, 2026, from approximately 05:20 to 12:55 \\(UTC\\), some customers experienced intermittent failures when using the Gmail Phish Alert Button add-in to report suspicious emails, receiving the \"An internal error has occurred\" message instead of a successful report confirmation. The issue affected customers across all regions, with the highest error volumes concentrated in the EU-West and US-East service regions.\n\nThis issue was caused by a disruption in a caching service on Google's Apps Script platform, which the add-in relies on to store per-user session configuration. The service began intermittently returning empty results on read while still accepting new writes, causing affected report submissions to fail. KnowBe4 engaged Google directly and filed a formal bug report between approximately 12:19 and 12:39 \\(UTC\\). Error volume declined sharply shortly after as the vendor-side condition cleared, and the Gmail Phish Alert Button returned to normal performance by approximately 12:55 \\(UTC\\).\n\nTo prevent this type of issue in the future, we are continuing to work with Google to track the resolution of the underlying caching defect. No data loss occurred as a result of this issue.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-08-21T14:57:08.327Z",
"resolved_inferred": false,
"started_at": "2026-08-21T08:52:22.748Z",
"state": "postmortem",
"title": "PAB | Error Reporting Phish using Gmail Phish Alert Button",
"updated_at": "2026-08-26T15:55:50.007Z",
"url": "https://stspg.io/97cd088qr41p"
},
{
"body": "This incident has been resolved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-08-18T20:14:25.041Z",
"resolved_inferred": false,
"started_at": "2026-08-18T14:29:27.958Z",
"state": "resolved",
"title": "PAB sending a Non-Delivery Report after an Email is Reported (US Only)",
"updated_at": "2026-08-18T20:14:25.053Z",
"url": "https://stspg.io/p1k0h84nj514"
},
{
"body": "Upon further investigation we have determined that the errors affecting Google PhishRIP queries were related to individual issues on Google accounts.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-08-06T21:56:23.335Z",
"resolved_inferred": false,
"started_at": "2026-08-05T19:29:46.279Z",
"state": "resolved",
"title": "Google Workspace PhishRIP errors",
"updated_at": "2026-08-06T21:56:23.352Z",
"url": "https://stspg.io/pbn7ggjs3kgd"
},
{
"body": "This incident has been resolved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-08-05T12:17:30.836Z",
"resolved_inferred": false,
"started_at": "2026-08-04T19:19:15.607Z",
"state": "resolved",
"title": "SAT Account Settings not Loading",
"updated_at": "2026-08-05T12:17:30.853Z",
"url": "https://stspg.io/sdc4j1sg5f6l"
},
{
"body": "## **Summary**\n\nOn July 28, 2026, an update to the PhishRIP service caused some PhishRIP search queries to match more broadly than they should have. Benign emails that didn't meet the configured sender criteria were quarantined across some customer accounts.\n\nAutomated monitoring and customer support tickets surfaced the problem quickly. Engineering reverted the change within approximately 25 minutes of declaring the incident and stopped the unintended quarantining. Engineering then manually restored the incorrectly quarantined emails to customers\u2019 inboxes.\n\n## **Root cause**\n\nOn July 28, 2026, we deployed a maintenance update to the query service behind PhishRIP. This update was designed to improve the handling of invalid sender addresses by allowing the system to fall back to searching only the domain portion of the address.\n\nHowever, existing validation logic used a regular expression that required an \"@\" symbol in the sender field. When a query used a bare domain, such as [example.com](http://example.com) with no username or \"@\" symbol, the regex flagged the domain as invalid.\n\nThis conflict between the existing validation and the domain-search update caused the system to drop sender criteria entirely. The query ran against inboxes using only its remaining parameters, matching and quarantining emails well outside the intended sender.\n\n## **Timeline**\n\nAll times are in UTC.\n\n| **Time** | **Event** |\n| --- | --- |\n| 19:36 | Our automated monitoring detected anomalous query execution patterns and alerted our engineers, who immediately began investigating. |\n| 19:51 | The issue was escalated to a high-severity production incident, and a response team was assembled. |\n| 19:52 | Our support team escalated the first customer tickets that reported unexpected quarantines, corroborating the monitoring alert. |\n| 19:54 | The root cause was traced to a recent queries API deployment, and engineering began reverting the change. |\n| 20:00 | Incident communications were posted to our status page, with PhishRIP marked as Degraded. |\n| 20:16 | The deployment was fully reverted. |\n| 20:35 | Production checks confirmed that subject and domain queries were again matching only their intended targets, and that no further messages were being quarantined in error. Our status page was updated accordingly. |\n| 20:39  | Engineering began restoring the incorrectly quarantined emails. |\n| 03:49 \\(next day\\) | Restorations were completed for all customer query sets logged through our support team. |\n\n## **Mitigation and remediation**\n\n**Code revert.** We weighed halting active PhishRIP operations mid-run against reverting the code. Halting the background tasks would have disrupted legitimate security workflows for every tenant, so reverting was the better option. The code revert was deployed and verified in production within approximately 40 minutes of escalation.\n\n**Email restoration.** To avoid releasing genuinely malicious emails back into customer environments, we manually restored them rather than in bulk. Engineering used internal admin tools to restore emails in batches, and worked with Support to verify each account ID, query ID, and receive customer approval first. All reported accounts were restored overnight.\n\n## **Preventative measures**\n\n* **Regression coverage**: We are adding unit and integration tests for domain-only, non-standard, and bare-string sender queries to catch these input cases before deployment.\n* **Query guardrails**: We are adding a check in the query engine that blocks queries from running if their core filtering parameters are dropped during evaluation.\n\n## **Conclusion**\n\nA single validation rule caused this: a regex that assumed every sender field contained an `@`. When queries used a bare domain, that assumption dropped the sender filter and the queries matched far more mail than intended. The fix was straightforward once we found it, and the revert stopped the damage within about 25 minutes of declaring the incident.\n\nRestoring quarantined mail was a harder process: we couldn't restore emails in bulk without risking genuinely malicious emails going back into inboxes. Our team chose to proceed carefully, working account by account with Support throughout the night.\n\nTwo preventative measures would have caught this earlier: a test covering domain-only sender queries, and a guardrail that refuses to run a query when its sender filter has been stripped. Both are now planned for. The immediate issue is resolved, and all reported accounts have been restored.\n\n## **Glossary**\n\n* **PhishRIP:** A PhishER feature that lets administrators search for and quarantine matching phishing emails across every user inbox in an organization.\n* **Regex \\(regular expression\\):** A pattern used in software to search text and validate input, such as checking that an email address is formatted correctly.\n* **Code revert:** Backing out a software update to return to the previous stable version.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-29T18:58:12.738Z",
"resolved_inferred": false,
"started_at": "2026-07-28T20:00:15.608Z",
"state": "postmortem",
"title": "Unexpected Emails being quarantined by PhishRIP",
"updated_at": "2026-07-30T16:53:16.577Z",
"url": "https://stspg.io/xn4ddvjkjgm7"
},
{
"body": "This incident has been resolved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-09-11T15:04:39.863Z",
"resolved_inferred": false,
"started_at": "2026-07-17T13:18:28.786Z",
"state": "resolved",
"title": "Protect blank purchasing page on store.knowbe4.com",
"updated_at": "2026-09-11T15:04:39.880Z",
"url": "https://stspg.io/cddcxl6t1s55"
},
{
"body": "On Thursday, July 16, 2026, from approximately 19:57 to 21:12 \\(UTC\\), some customers experienced difficulty logging in to the Defend console, seeing a **Find Out More** screen instead of console access.\n\nThis issue was caused by a synchronization issue between Salesforce and our internal licensing systems, which incorrectly set some customer license counts to zero. As a result, affected customers were unable to log in. To resolve this issue, our team applied a temporary license override to restore access for affected customers while correcting the underlying licensing records in Salesforce. Once corrected, the systems automatically synchronized the accurate values, and the Defend console access returned to normal performance by approximately 21:12 \\(UTC\\).\n\nTo prevent this type of issue in the future, we are implementing improvements to the Salesforce synchronization process to ensure accurate license counts.\n\nNo data loss occurred as a result of this issue.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-16T21:20:20.727Z",
"resolved_inferred": false,
"started_at": "2026-07-16T19:48:18.636Z",
"state": "postmortem",
"title": "Defend - Redirect Upon Login",
"updated_at": "2026-08-20T21:32:25.964Z",
"url": "https://stspg.io/btwk3mml7w2v"
},
{
"body": "This incident has been resolved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-16T14:04:21.572Z",
"resolved_inferred": false,
"started_at": "2026-07-15T16:46:33.000Z",
"state": "resolved",
"title": "Protect blank purchasing page on store.knowbe4.com",
"updated_at": "2026-07-16T14:04:21.585Z",
"url": "https://stspg.io/gn0q3mjn62z8"
},
{
"body": "An upstream feature flag service issue degraded model routing in collaboration-inference, causing PhishML evaluations to return 0/0/0 default errors. The third party deployed a fix, and a code update was deployed to production adding static fallback routing.\nAll services have recovered and are operating normally.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-10T14:30:00.000Z",
"resolved_inferred": false,
"started_at": "2026-07-10T14:30:00.000Z",
"state": "resolved",
"title": "PhishML Degradation",
"updated_at": "2026-07-22T21:01:10.823Z",
"url": "https://stspg.io/fnqbqws7cr8v"
},
{
"body": "## **Executive Summary**\n\nOn July 1, 2026, the KnowBe4 Security Awareness Training \\(KSAT\\) platform experienced a period of degraded performance resulting in intermittent login errors and high latency for users on our United States \\(US\\) instance. The issue was initiated following a routine platform deployment that introduced an unoptimized database query. This query placed an excessive operational load on our primary database reader cluster, causing database sessions to saturate and subsequent login requests to queue up.\n\nEngineering teams promptly identified the degradation, reverted the deployment, and systematically cleared the backlogged database sessions to restore optimal performance. The issue did not affect data integrity or security, and service was fully stabilized.\n\n## **Technical Root Cause**\n\nThe root cause was determined to be a newly introduced query within a standard application update. Upon deployment, this specific query pattern bypassed optimal indexing strategies, resulting in full table scans and highly extended execution times on the database reader infrastructure.\n\nAs a high volume of authentication and application requests arrived concurrently, the database reader quickly exhausted its available connection pool due to these long-running, unmitigated database sessions. This resource starvation immediately manifested as severe application latency and intermittent timeouts during the user authentication process.\n\n## **Timeline of Events**\n\nThe incident began in the mid-morning hours and progressed through identification, remediation, and verification stages over a period of approximately 152 minutes.\n\nAn internal high-severity incident response group was established immediately following automated monitoring alerts indicating that health checks targeting our application programming interface \\(API\\) routing layer were failing from regional cloud monitoring nodes. This behavior was confirmed by concurrent engineering analysis of user HTTP Archive files, which demonstrated severe latency spikes specifically isolated to the authentication endpoints.\n\nWithin three minutes of establishing the response team, cross-referencing recent system changes pointed to a recent application deployment as the primary catalyst. Engineers immediately initiated the rollback process, drafting and approving a revert modification to extract the problematic code from the deployment pipeline.\n\nThe deployment of the reverted codebase to the US production cluster commenced shortly thereafter. While the deployment processed over the subsequent twenty-five minutes, technical personnel prepared direct database interventions to clear the residual system strain. Once the stable code version was completely active across the fleet, engineers began systematically terminating the lingering, long-running database sessions that had been spawned by the unoptimized query. To ensure a pristine state, the database reader infrastructure was cycled twice.\n\nFollowing these administrative infrastructure restarts, operational telemetry showed the database reader load dropping significantly to a healthy baseline of approximately thirty-six percent. System performance normalized, and administrative logging verified that authentication requests were processing within standard latency thresholds. After monitoring the environment to confirm sustained stability, engineers officially marked the incident as mitigated, later shifting the status to fully resolved following an extended window of zero performance spikes and complete passes on all automated sanity test suites.\n\n## **Mitigation**\n\nTo alleviate the immediate infrastructure distress and restore user access, the engineering team executed a multi-phased mitigation strategy:\n\n* **Codebase Rollback:** The changes introduced in the recent deployment were immediately isolated, reverted, and redeployed to production to prevent any further generation of the unoptimized query.\n* **Database Session Termination:** Internal engineering tools were utilized to explicitly terminate active, long-running database queries that were blocking the connection pools.\n* **Infrastructure Cycling:** The database reader instances were restarted twice in succession to flush out stale memory allocations and guarantee that all orphaned database sessions were permanently cleared.\n\n## **Preventative Measures**\n\nTo prevent a recurrence of this specific issue and mitigate the impact of similar query-based database bottlenecks in the future, KnowBe4 is implementing the following actions:\n\n* **Enhanced Query Linting and Analysis:** Integrate automated query execution plan analysis into our continuous integration and continuous deployment pipelines to flag unindexed or high-cost queries before they reach production environments.\n* **Database Connection Pooling Guardrails:** Adjust database timeouts and implement aggressive circuit-breaker thresholds for user authentication paths to prevent single, long-running query patterns from exhausting the entire connection pool.\n* **Load Shedding Policies:** Implement strict application-level timeouts on read-heavy database calls to ensure they fail gracefully rather than degrading the overall availability of the core login workflows.\n\n## **Conclusion**\n\nWe sincerely apologize for the inconvenience and friction this performance degradation caused our customers and partners. KnowBe4 is dedicated to maintaining high availability and reliability across our product suites. By refining our pre-deployment automated query validation and bolstering our database connection resiliency, we are actively working to ensure the continuous, seamless operation of the KSAT platform.\n\n## **Glossary of Technical Terms**\n\n* **API \\(Application Programming Interface\\):** A set of protocols that allows different software applications to communicate with one another. In this context, it routes authentication requests from the user interface to the backend servers.\n* **Database Reader:** A dedicated database instance or cluster responsible for handling read-only queries \\(such as fetching user profiles or validating login configurations\\), separating this traffic from write operations to optimize performance.\n* **HAR \\(HTTP Archive\\) File:** A JSON-formatted log file that records a web browser's interaction with a website, used by engineers to diagnose performance and network latency issues.\n* **Latency:** The time delay or duration it takes for a data packet or request to travel from its source to its destination and return a response.\n* **Sanity Suite:** A collection of automated tests executed against a deployment environment to quickly verify that the core functionality of an application is working correctly.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-01T18:01:17.670Z",
"resolved_inferred": false,
"started_at": "2026-07-01T14:40:41.868Z",
"state": "postmortem",
"title": "KSAT Intermittent Login Issues - US",
"updated_at": "2026-07-08T18:52:50.209Z",
"url": "https://stspg.io/k9xxcyp7v05w"
},
{
"body": "#### Summary\n\nOn June 30, 2026, customers using the Defend US service experienced delays in email processing. The incident began at approximately 14:00 UTC and was fully resolved by 19:45 UTC.\n\n\u200c\n\n#### What Happened\n\nA scheduled maintenance operation began in the early morning of June 30. As this operation progressed, it placed an unexpectedly high load on our infrastructure, which caused email processing to slow down across our US service.\n\n\u200c\n\nOur team identified the issue and declared an incident at 16:00 UTC. Steps were taken to reduce the load and restore normal processing speeds, including pausing non-essential background activity and engaging our infrastructure provider for additional support.\n\n\u200c\n\nBy 17:45 UTC, email delivery delays had been fully resolved. Email analysis continued to recover and was back to normal by 19:45 UTC.\n\n\u200c\n\n#### Customer Impact\n\nEmail delivery \\(SMTP customers\\): Emails were delayed in transit by up to 30 minutes between approximately 14:00 UTC and 17:45 UTC. All emails were delivered; no messages were lost.\n\n\u200c\n\nEmail analysis \\(Microsoft 365 / Graph API customers\\): Email analysis was delayed by an average of 30 minutes before emails were analysed, between approximately 14:00 UTC and 19:45 UTC. Email delivery to end users was not affected.\n\n\u200c\n\n#### What We Are Doing\n\nWe have rescheduled the maintenance operation that triggered this incident to run during an overnight, low-traffic window, giving it sufficient time to complete without affecting the live service.\n\n\u200c\n\nWe are also investing in infrastructure improvements to better isolate maintenance operations from customer-facing workloads, so that future maintenance cannot affect email processing in this way.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-06-30T19:46:45.179Z",
"resolved_inferred": false,
"started_at": "2026-06-30T16:14:53.577Z",
"state": "postmortem",
"title": "Defend - Increased Email Latency (US Only)",
"updated_at": "2026-07-03T15:24:21.577Z",
"url": "https://stspg.io/5s14r8s5mq15"
},
{
"body": "On Tuesday, June 30, 2026, from approximately 07:40 to 19:15 \\(UTC\\), customers experienced incorrect results from PhishER's PhishML scoring. Affected emails received a PML:BYPASSED tag instead of a legitimate PhishML classification, and confidence scores were missing from impacted messages. Rules and actions that depend on PhishML results also did not activate.\n\nThis issue was caused by a code refactor introduced approximately two weeks earlier. This refactor introduced a faulty update that omitted essential drivers required for PhishML scoring to run. However, the issue remained dormant until another update triggered a new PhishML model deployment, which caused the scoring issue to emerge. To resolve this issue, our team rolled back to the last stable deployment and added more capacity to process the resulting backlog of email evaluations. PhishER's PhishML scoring returned to normal performance by 19:15 \\(UTC\\).\n\nTo prevent this type of issue in the future, we have improved health checks by introducing a new endpoint for smoke testing new models before deployment.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-06-30T19:11:55.790Z",
"resolved_inferred": false,
"started_at": "2026-06-30T14:18:11.378Z",
"state": "postmortem",
"title": "PhishML Evaluations Causing PML:BYPASSED Tags to Apply",
"updated_at": "2026-07-22T13:20:49.661Z",
"url": "https://stspg.io/xt1mc4pk5tdn"
},
{
"body": "From Tuesday, June 23, 2026, at approximately 10:03 \\(UTC\\) to Thursday, June 25, 2026, at approximately 11:54 \\(UTC\\), some US and EU customers experienced intermittent 502 errors when logging in to KCM GRC.\n\nThis issue was caused by network traffic attempting to connect to invalid multilevel subdomains, which overwhelmed the cache serving KCM GRC and resulted in login errors. Our team initially updated the configuration of our content delivery network, which temporarily resolved the errors, but they returned later that day. After further investigation, we identified the caching issue as the root cause and deployed an infrastructure-level fix to prevent multi-level subdomain traffic from affecting the cache. KCM GRC returned to normal performance by 11:54 \\(UTC\\) on June 25, 2026.\n\nTo prevent this type of issue in the future, we are evaluating additional protections to guard against similar traffic that could affect the cache.\n\nNo data loss occurred as a result of this issue.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-06-25T13:10:56.444Z",
"resolved_inferred": false,
"started_at": "2026-06-24T15:48:26.151Z",
"state": "postmortem",
"title": "KCM GRC - 500 Errors Upon Login",
"updated_at": "2026-08-04T14:50:48.979Z",
"url": "https://stspg.io/blhxy7cq7jsc"
},
{
"body": "From Tuesday, June 16, 2026, at approximately 17:34 \\(UTC\\), to Thursday, June 18, 2026, at approximately 22:41 \\(UTC\\), some customers experienced intermittent unavailability of the **Phishing Security Test Reports** page in the KnowBe4 console.\n\nThis issue was caused by a code change that introduced a conflict between two methods for processing phishing campaign data. As a result, phishing campaigns still using legacy phishing categories were unable to load the **Phishing Security Test Reports** page. To resolve this issue, we updated the code to process campaigns correctly under both classification systems, and the **Phishing Security Test Reports** page returned to normal performance by June 18, 2026, at 22:41 \\(UTC\\).\n\nNo data loss occurred as a result of this issue.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-06-19T14:01:57.153Z",
"resolved_inferred": false,
"started_at": "2026-06-18T21:14:38.770Z",
"state": "postmortem",
"title": "Phishing test report tab unavailable",
"updated_at": "2026-07-07T19:45:10.291Z",
"url": "https://stspg.io/4gtp2g55nxc4"
},
{
"body": "From Wednesday, June 17, 2026, at approximately 19:00 \\(UTC\\), to Thursday, June 18, 2026, at approximately 18:20 \\(UTC\\), some customers were unable to view results in the KnowBe4 Security Center\u2019s Human Risk Management widget.\n\n\u200c\n\nThis issue was caused by a recent infrastructure migration that left an internal service connection pointing to an outdated endpoint. Though the underlying data processing was unaffected, the Human Risk Management widget could not retrieve data or display results. To resolve this issue, our team updated and reapplied the connection configuration, and the KnowBe4 Security Center returned to normal performance by approximately 18:20 \\(UTC\\) on June 18, 2026.\n\nTo prevent this type of issue in the future, we are improving our migration process to ensure that all endpoints are valid after a migration. We are also strengthening known-good rollback procedures in cases where a migration cannot be completed as planned.\n\nNo data loss occurred as a result of this issue.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-06-18T19:12:35.819Z",
"resolved_inferred": false,
"started_at": "2026-06-18T15:13:09.435Z",
"state": "postmortem",
"title": "KnowBe4 Security Center (KSC) | Human Risk Managment Widget Showing No Results",
"updated_at": "2026-09-08T15:45:59.358Z",
"url": "https://stspg.io/m5t4dpw6q2j1"
},
{
"body": "From Tuesday, June 16, 2026, at approximately 22:23 \\(UTC\\), to Wednesday, June 17, 2026, at approximately 20:19 \\(UTC\\), some customers experienced inaccurate user counts in KSAT reporting. Downstream features that rely on this data were also affected, including AIDA Orchestration, Risk Score, and ModStore recommendations.\n\nThis issue was caused by an incomplete rebuild of the users' data table. A data storage policy removed historical data earlier than intended, so when a full table rebuild ran, some older data was unavailable, causing user counts to drop. To resolve this issue, our team performed a full data migration sync on the users' table and retriggered the dependent data pipelines to update downstream systems. We then tested and confirmed that user counts were returning correct results, and the KSAT console returned to normal performance by June 17, 2026, at 20:19 \\(UTC\\).\n\nTo prevent this type of issue in the future, we are correcting the underlying data storage policy to ensure complete historical data is retained for the full rebuild process.\n\nNo data loss occurred as a result of this issue.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-06-17T20:19:17.108Z",
"resolved_inferred": false,
"started_at": "2026-06-16T22:23:16.540Z",
"state": "postmortem",
"title": "Data and User Inconsistencies in Reporting",
"updated_at": "2026-08-21T18:01:36.229Z",
"url": "https://stspg.io/4tkf5xjd4z79"
},
{
"body": "On Monday, June 15, 2026, from approximately 15:30 to 16:34 \\(UTC\\), some customers experienced processing delays with phishing and training campaigns in the KnowBe4 console.\n\nThis issue was caused by an update that introduced an incompatible software version. This software prevented our background processing service from completing queued tasks. As a result, training enrollments, notification sending, and SmartGgroup enrollments were delayed. To resolve this issue, our team deployed a fix that reverted the affected dependency and restored normal job processing. Once the fix was deployed, the affected queues cleared, and the KnowBe4 console returned to normal performance by 16:34 \\(UTC\\).\n\nTo prevent this issue in the future, we have implemented additional testing layers for similar deployments.\n\nNo data loss occurred as a result of this issue.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-06-15T16:56:34.104Z",
"resolved_inferred": false,
"started_at": "2026-06-15T16:56:34.054Z",
"state": "postmortem",
"title": "Latency Issues",
"updated_at": "2026-08-20T19:41:18.677Z",
"url": "https://stspg.io/k39rjhspld1r"
},
{
"body": "On Wednesday, June 3, 2026, from approximately 14:49 to 16:09 \\(UTC\\), customers experienced errors when accessing the **User Details** page in the KSAT console.\n\nThis issue was caused by an update that introduced missing fields, which prevented the **User Details** page from loading. To resolve this issue, our team updated the console again to restore the missing fields, and the KSAT console returned to normal performance by approximately 16:09 \\(UTC\\). To prevent similar issues in the future, additional automated tests were added.\n\nNo data loss occurred as a result of this issue.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-06-03T18:14:47.592Z",
"resolved_inferred": false,
"started_at": "2026-06-03T15:07:44.987Z",
"state": "postmortem",
"title": "KSAT - Widespread User Profile Access Issues",
"updated_at": "2026-07-15T15:23:51.961Z",
"url": "https://stspg.io/bxsdm75mg6kq"
},
{
"body": "On Thursday, May 28, 2026, from approximately 18:07 to 18:57 \\(UTC\\), some customers experienced issues uploading custom content to the new ModStore in KnowBe4 Security Awareness Training \\(KSAT\\).\n\nThis issue was caused by an update to the new ModStore that disabled the language selection field in the \u201cAdd Translation\u201d upload step, preventing customers from completing that required part of the upload process. To resolve this issue, our engineering team identified the recent deployment responsible for the defect, corrected it, and deployed the update to production. KSAT returned to normal performance by 18:57 \\(UTC\\).\n\nTo prevent this type of issue in the future, we are improving test coverage for the affected upload flow and adding automated checks to catch similar issues before they reach production.\n\nNo data loss occurred as a result of this issue.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-05-28T19:13:16.141Z",
"resolved_inferred": false,
"started_at": "2026-05-28T18:17:52.922Z",
"state": "postmortem",
"title": "Unable to upload custom content to the new Modstore",
"updated_at": "2026-07-07T19:43:14.827Z",
"url": "https://stspg.io/gxmg8ggjlr0p"
},
{
"body": "This incident has been resolved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-05-27T15:44:11.607Z",
"resolved_inferred": false,
"started_at": "2026-05-27T10:05:45.801Z",
"state": "resolved",
"title": "Protect -  Intermittent Issues with Sending Emails (UK Only)",
"updated_at": "2026-05-27T15:44:11.623Z",
"url": "https://stspg.io/g73wyjkt9zm9"
},
{
"body": "# **Summary**\n\nOn Wednesday, May 21, 2026, customers using the new ModStore experience within KnowBe4 Security Awareness Training \\(KSAT\\) were unable to access the ModStore and received HTTP 500 errors. The disruption affected all regional instances and lasted approximately 23 minutes, from 13:17 to 13:40 \\(UTC\\).\n\nThis issue was caused by a defective code change in a routine ModStore deployment. The deployment itself completed successfully, but the change caused the application to return server errors when customers attempted to load the new ModStore. Our engineering team identified the faulty deployment within minutes, rolled it back, and confirmed full restoration of service at 13:40 \\(UTC\\).\n\nOnly the new ModStore experience was affected. Customers using the classic ModStore, and all other KSAT functionality \u2014 including training assignments, campaigns, and reporting \u2014 remained fully operational throughout.   \n  \nThere was no data loss as a result of this incident.\n\n# **What Happened**\n\n## **New ModStore \\(KSAT\\)**\n\nAt 13:08 \\(UTC\\) on May 21, 2026, a routine deployment of the new ModStore application began and was completed successfully across all production environments at 13:17 \\(UTC\\). The release included changes related to integrating the new ModStore natively into the KnowBe4 platform. Immediately following the deployment, the new ModStore began returning HTTP 500 errors to customers attempting to access it.\n\nAt 13:28 \\(UTC\\), internal reports of the errors reached the engineering team, and by 13:30 \\(UTC\\) engineers had correlated the failures with the deployment. The decision to roll back was made at 13:31 \\(UTC\\). At 13:32 \\(UTC\\), with multiple customer support tickets confirming customer-facing impact, a high-severity incident was declared, and on-call engineers were contacted.\n\nThe team quickly confirmed the scope: only the new ModStore experience, enabled for approximately half of customers plus those who had opted in, was affected, across all regional instances. The classic ModStore and all other KSAT functionality remained available. A rollback to the previous stable version was initiated at 13:35 \\(UTC\\) and completed at 13:40 \\(UTC\\), at which point access to the new ModStore was fully restored. Our public status page was updated at 13:40 \\(UTC\\) and moved to \u201cMonitoring\u201d one minute later.\n\nDuring the post-restoration investigation, an engineer noted that a separate infrastructure-as-code deployment earlier that morning had unexpectedly altered a storage Cross-Origin Resource Sharing \\(CORS\\) configuration used by the ModStore. This change was investigated as a potential contributor and ruled out \u2014 it was not related to the 500 errors. The investigation did, however, surface that two separate deployment pipelines were both managing the same infrastructure configuration and silently overwriting each other\u2019s changes. This conflict was remediated during the incident window by assigning the configuration a single owning pipeline. After continued monitoring confirmed stability, the incident was resolved internally at 14:52 \\(UTC\\), and the public status page was updated to \u201cResolved\u201d at 17:20 \\(UTC\\).\n\n# **Root Cause Analysis**\n\nThe root cause of this incident was a defective code change included in the 13:17 \\(UTC\\) deployment of the new ModStore application. Once deployed, the change caused the application to fail to serve customer requests, returning HTTP 500 errors to all users of the new ModStore experience across all regions. The defect was not detected during pre-deployment testing, so the deployment proceeded to production as usual.\n\nBecause the failure began at the moment the deployment completed and affected all regions simultaneously, engineers were able to identify the deployment as the trigger within minutes and end customer impact by rolling back to the previous version. The faulty change was withheld for rework and additional validation before any reintroduction.\n\nA secondary issue was identified during the investigation: an unrelated infrastructure configuration change that morning initially appeared connected because of its timing. It was ruled out as a cause, but the investigation revealed that two deployment pipelines shared ownership of the same infrastructure configuration and were overwriting each other. While this did not cause the incident, it added noise to the diagnosis. Because it represented a latent risk, it was also corrected the same day.\n\n## **Detailed Timeline \\(UTC\\)**\n\n| **Time \\(UTC\\)** | **Event** |\n| --- | --- |\n| **13:08** | A routine deployment to the new ModStore application begins. |\n| **13:17** | The deployment completes across all production environments; HTTP 500 errors begin for the new ModStore experience. |\n| **13:28** | Internal reports of errors loading the new ModStore reach the engineering team. |\n| **13:30** | Engineers correlate the HTTP 500 errors with the earlier deployment. |\n| **13:31** | Decision made to roll back the deployment. |\n| **13:32** | A high-severity incident is declared; on-call engineers are paged. Customer support tickets confirm customer-facing impact. |\n| **13:33 \u2013 13:34** | Impact confirmed to be limited to the new ModStore experience; the classic ModStore and all other KSAT functionality are confirmed unaffected. All regional instances are affected. |\n| **13:35 \u2013 13:38** | The revert is prepared, and the rollback deployment begins. |\n| **13:40** | Rollback completes, and access to the new ModStore is restored. The public status page is updated to \u201cIdentified.\u201d |\n| **13:41** | The public status page is updated to \u201cMonitoring.\u201d |\n| **13:58 \u2013 14:11** | An unrelated same-morning infrastructure configuration change is investigated and ruled out as a cause. A conflicting ownership issue between two deployment pipelines that manage the same configuration is identified and remediated. |\n| **14:52** | After continued monitoring confirms stability, the incident is marked resolved internally \\(74 minutes after the incident was declared\\). |\n| **17:20** | The public status page is updated to \u201cResolved.\u201d |\n\n\u200c\n\n# **Findings and Mitigations**\n\n## **1. A defective change reached production**\n\nA code change included in a routine deployment of the new ModStore caused the application to return server errors in production. The defect was not caught by the automated tests that run before a release is promoted, indicating a gap in pre-deployment test coverage for this failure mode.\n\n**Mitigations:**\n\n* The deployment was rolled back within 23 minutes of impact beginning, immediately restoring service.\n* The faulty change was withheld from redeployment pending rework and additional validation.\n* Pre-deployment test coverage for the new ModStore is being expanded to cover the failure mode seen in this incident.\n\n## **2. Detection relied on human reports rather than automated alerting**\n\nImpact began at 13:17 \\(UTC\\), but the engineering team was first engaged through internal reports at 13:28 \\(UTC\\) and customer support tickets shortly after, rather than by an automated alert on the application\u2019s error rate. Automated post-deployment checks that ran after the release did not halt the rollout or page the team.\n\n**Mitigations:**\n\n* Error-rate monitoring and automated alerting for the new ModStore are being strengthened so that a spike in server errors immediately after a deployment pages the on-call team directly.\n* Post-deployment verification is being reviewed so that failing checks more decisively block or flag a release.\n\n## **3. Two deployment pipelines managed the same infrastructure configuration**\n\nThe application\u2019s deployment pipeline and a separate infrastructure-as-code pipeline both defined the same storage CORS configuration, and each deployment silently overwrote the other\u2019s settings. This conflict was not the cause of the incident, but it initially complicated diagnosis and represented an ongoing risk of unintended configuration changes.\n\n**Mitigations:**\n\n* The duplicated configuration was removed from the application pipeline and its state references were cleaned up, making the infrastructure pipeline the single owner of that configuration.\n* A review is underway to identify any other resources with shared ownership across pipelines.\n\n# **Customer Impact**\n\n## **KSAT \u2013 New ModStore \\(all regions\\)**\n\nFrom 13:17 to 13:40 \\(UTC\\), customers who had the new ModStore experience enabled \\(approximately half of customers, plus those who opted in\\) received HTTP 500 errors when attempting to access the ModStore and were unable to browse or add training content during that window. All regional instances were affected equally.\n\nCustomers using the classic ModStore were not affected. All other KSAT functionality \u2014 including active training campaigns, phishing simulations, user enrollments, and reporting \u2014 operated normally throughout the incident. No customer action was or is required, and no data was lost or altered.\n\n# **Preventive Measures**\n\n* **Expanded pre-deployment testing \u2014** Test coverage for the new ModStore is being extended to catch the class of defect that caused this incident before a release reaches production.\n* **Automated error-rate alerting \u2014** Monitoring on the new ModStore is being strengthened so that elevated server-error rates following a deployment automatically notify on-call engineers, removing the dependency on human reports for detection.\n* **Stronger post-deployment verification \u2014** The automated checks that run immediately after a release are being reviewed so that failures block or escalate a rollout rather than passing silently.\n* **Single ownership of infrastructure configuration \u2014** Shared infrastructure settings are being consolidated under a single owning pipeline to prevent conflicting automated changes, and an audit is underway to identify any remaining overlaps.\n\n# **Conclusion**\n\nOn May 21, 2026, a defective code change during a routine deployment made the new ModStore experience unavailable for approximately 23 minutes, resulting in server errors for affected customers across all regions. We recognize that the ModStore is central to how administrators build their training programs, and we apologize for the disruption.\n\nThe response demonstrated the value of fast rollback as a recovery path: the faulty release was identified and reverted within minutes of the first reports, the scope was accurately confirmed early, and no data was lost. The incident also highlighted clear opportunities to improve: stronger pre-deployment testing, automated error-rate detection, and cleaner ownership of infrastructure configuration. These are being addressed through the preventive measures above.\n\nWe are committed to the reliability of the new ModStore experience as it rolls out to all customers, to ensuring that issues of this kind are caught before they reach production, and to detecting and resolving issues automatically if they do reach production.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-05-21T17:20:20.582Z",
"resolved_inferred": false,
"started_at": "2026-05-21T13:40:58.947Z",
"state": "postmortem",
"title": "New Modstore 500 Errors",
"updated_at": "2026-07-07T15:35:02.251Z",
"url": "https://stspg.io/nxp2k09gjg8s"
},
{
"body": "From Monday, May 18, 2026, at approximately 14:00 \\(UTC\\), to Tuesday, May 19, 2026, at approximately 15:00 \\(UTC\\), users experienced timeouts when attempting to use the Microsoft Ribbon Phish Alert Button \\(PAB\\) in Classic Outlook.\n\nThis issue was caused by an update to the PAB's authentication process that was incompatible with the Classic Outlook environment. The update used a loading method that Classic Outlook does not support, which caused the PAB to time out before it could connect to KnowBe4's servers.\u00a0\n\nTo resolve this issue, we updated how the authentication library is packaged with the PAB to ensure compatibility with Classic Outlook. The Microsoft Ribbon PAB returned to normal performance by approximately 15:00 \\(UTC\\) on May 19, 2026.\n\nTo prevent this type of issue in the future, we are updating our processes for compatibility testing and improving monitoring.\n\nNo data loss occurred as a result of this issue.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-05-19T18:04:42.563Z",
"resolved_inferred": false,
"started_at": "2026-05-19T14:16:44.172Z",
"state": "postmortem",
"title": "Microsoft Ribbon PAB Operation Timeout",
"updated_at": "2026-06-25T18:39:16.717Z",
"url": "https://stspg.io/fltk6zvgb6j9"
},
{
"body": "Our team has identified and resolved the issue causing latency with the KSAT console",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-05-18T21:40:44.113Z",
"resolved_inferred": false,
"started_at": "2026-05-18T17:28:51.875Z",
"state": "resolved",
"title": "KSAT Latency Issues",
"updated_at": "2026-05-18T21:40:44.128Z",
"url": "https://stspg.io/r0f7y6z1nz4w"
},
{
"body": "On Wednesday, May 13, 2026, from approximately 18:00 until 18:45 \\(UTC\\), some users experienced issues with accessing the KnowBe4 Academy.  \n\nThis incident was caused by a service disruption with an external service provider. The provider restored service and the Academy returned to normal performance by 18:45 \\(UTC\\).\n\nNo data loss occurred as a result of this issue.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-05-14T13:07:32.769Z",
"resolved_inferred": false,
"started_at": "2026-05-13T18:09:24.097Z",
"state": "postmortem",
"title": "KnowBe4 Academy - Access Issues",
"updated_at": "2026-05-14T19:08:03.725Z",
"url": "https://stspg.io/3t1tx2yfyldc"
},
{
"body": "On Tuesday, May 12, 2026, from approximately 10:12 until 11:43 \\(UTC\\), some US customers experienced issues while attempting to upload files or process data within the Workspace console.\n\nThis incident was caused by a networking disruption that affected a backend messaging service, preventing it from properly syncing data. While file downloads remained available, new file uploads and certain background tasks were blocked. To resolve this issue, we temporarily paused new network traffic and restored the messaging service. The Workspace console returned to normal performance by 11:43 \\(UTC\\).\n\nTo prevent this type of issue in the future, we are reviewing our messaging service configuration and improving network monitoring to more quickly identify synchronization issues.\n\nNo data loss occurred as a result of this issue.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-05-12T15:16:01.448Z",
"resolved_inferred": false,
"started_at": "2026-05-12T10:52:31.827Z",
"state": "postmortem",
"title": "Unable to upload items on Workspace - US",
"updated_at": "2026-05-21T16:26:19.086Z",
"url": "https://stspg.io/3hg37wvr9vp1"
},
{
"body": "On Friday, May 8, 2026, from approximately 01:45 to 03:08 \\(UTC\\), some US customers experienced issues accessing the PhishER console and PasswordIQ.\n\nThis incident was caused by a localized infrastructure failure at our cloud service provider, which affected authentication traffic and prevented login verifications. To resolve this issue, we reprovisioned our authentication resources to healthy environments. After we rerouted login traffic to the healthy environments, PhishER and PasswordIQ returned to regular performance by 03:08 \\(UTC\\).\n\nNo data loss occurred as a result of this issue.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-05-08T03:54:19.715Z",
"resolved_inferred": false,
"started_at": "2026-05-08T02:05:27.831Z",
"state": "postmortem",
"title": "PhishER - Unable to Login (US)",
"updated_at": "2026-05-15T14:49:54.895Z",
"url": "https://stspg.io/tt7wn4r4953b"
},
{
"body": "We have identified an issue with customers signing into PhishER in the UK instance and implemented a fix. Customers should be able to login successfully.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-04-30T16:47:16.662Z",
"resolved_inferred": false,
"started_at": "2026-04-30T13:47:54.885Z",
"state": "resolved",
"title": "PhishER - Unable to Login (UK)",
"updated_at": "2026-04-30T16:47:16.685Z",
"url": "https://stspg.io/299x95lm5rgt"
},
{
"body": "We have identified an issue with customers pushing ADI files and training policy uploads to KSAT. Our team has implemented a fix and this has now been resolved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-04-27T17:19:13.187Z",
"resolved_inferred": false,
"started_at": "2026-04-27T13:40:10.226Z",
"state": "resolved",
"title": "ADI Sync and Training Policy Upload Errors (US Only)",
"updated_at": "2026-04-27T17:19:13.205Z",
"url": "https://stspg.io/r22tr3ls7739"
},
{
"body": "This incident has been resolved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-04-27T11:09:31.402Z",
"resolved_inferred": false,
"started_at": "2026-04-24T20:35:15.019Z",
"state": "resolved",
"title": "PhishER Login Issues (CA Only)",
"updated_at": "2026-04-27T11:09:31.419Z",
"url": "https://stspg.io/xjfp19kh4hpw"
},
{
"body": "This incident has been resolved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-04-27T11:11:56.636Z",
"resolved_inferred": false,
"started_at": "2026-04-21T20:10:43.647Z",
"state": "resolved",
"title": "Intermittent latency for Protect users - US Only",
"updated_at": "2026-04-27T11:11:56.654Z",
"url": "https://stspg.io/qwbv7fsxnvtr"
},
{
"body": "On Monday, April 20, 2026, from approximately 14:00 until 20:00 \\(UTC\\), some customers experienced \"Restricted\" errors when attempting to access the PhishER console in our UK and Canada instances.\n\nThis incident was caused by a configuration mismatch during a service deployment. A change in our infrastructure code inadvertently removed a required regional provider alias and upgraded a core distribution module to an unverified version. These combined factors caused regional resources to be resolved incorrectly, which prevented some users from accessing the console. Once we restored the configuration to the correct provider alias and module version, the PhishER console returned to normal performance by 20:00 \\(UTC\\).\n\nTo prevent this type of issue in the future, we have implemented stricter rules for software versioning and improved our testing process. We\u2019ve also updated our internal documentation and monitoring alerts to identify and respond to similar configuration errors more quickly.\n\nNo data loss occurred as a result of this issue.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-04-20T18:50:09.978Z",
"resolved_inferred": false,
"started_at": "2026-04-20T15:18:13.323Z",
"state": "postmortem",
"title": "PhishER Access Issue for UK/CA Instances",
"updated_at": "2026-06-30T19:32:31.238Z",
"url": "https://stspg.io/7rn2l0p65kxf"
},
{
"body": "On Tuesday, March 31, 2026, from approximately 15:00 \\(UTC\\) to 16:10 \\(UTC\\), some users experienced issues with the Defend Abuse Mailbox service.\n\nThis incident was caused by a service update that started sending duplicate phishing submission notifications to users. To resolve this issue, we corrected the service update code that caused the duplicate notifications, and the Defend Abuse Mailbox service returned to normal performance by 16:10 \\(UTC\\).\n\nTo prevent this type of issue in the future, we are reviewing and improving our service upgrade workflow.\n\nNo data loss occurred as a result of this issue.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-03-31T16:45:44.637Z",
"resolved_inferred": false,
"started_at": "2026-03-31T15:13:03.564Z",
"state": "postmortem",
"title": "Defend - Sending multiple abuse mailbox submission notifications",
"updated_at": "2026-04-01T21:44:21.931Z",
"url": "https://stspg.io/9q5h9qjrcc3n"
},
{
"body": "On Thursday, March 19, 2026, from approximately 11:56 \\(UTC\\) until 13:14 \\(UTC\\), some UK users experienced intermittent issues while attempting to log in to Prevent and Protect consoles.\n\nThis incident was caused by a configuration timeout in our session management system. To resolve this issue, we increased the timeout limit, and the Prevent and Protect consoles returned to normal performance by 13:14 \\(UTC\\).\n\nTo prevent this type of issue in the future, we have improved our timeout monitoring process and are verifying configuration consistency and optimization across our environments.\n\nNo data loss occurred as a result of this issue.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-03-20T12:18:18.117Z",
"resolved_inferred": false,
"started_at": "2026-03-19T12:33:34.732Z",
"state": "postmortem",
"title": "Prevent and Protect - Administration Panel Slow or Inaccessible",
"updated_at": "2026-03-26T13:42:26.471Z",
"url": "https://stspg.io/2zx5jj4c23x7"
},
{
"body": "From Tuesday, March 17, 2026, at approximately 16:00 \\(UTC\\), to Wednesday, March 18, 2026, at approximately 13:00 \\(UTC\\), some users experienced errors while submitting Workspace webforms.\n\nThis incident was caused by a configuration error during a domain migration. To resolve this issue, we corrected the domain configuration, and webform submissions returned to normal performance by 13:00 \\(UTC\\) on March 18, 2026.\n\nTo prevent this type of issue in the future, we are reviewing and updating our domain migration process. \n\nNo data loss occurred as a result of this issue.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "critical",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-03-18T14:28:44.365Z",
"resolved_inferred": false,
"started_at": "2026-03-18T10:33:54.018Z",
"state": "postmortem",
"title": "Webforms not loading",
"updated_at": "2026-03-24T21:05:07.022Z",
"url": "https://stspg.io/nml3hw6hvbrw"
},
{
"body": "On Tuesday, March 10, 2026, from approximately 15:22 until 20:35 \\(UTC\\), the EU and US instances of the PhishER console displayed an incorrect subject line for some messages reported with the Phish Alert Button \\(PAB\\).\n\nThis incident was caused by a configuration change to the PhishER message ingester service, which unexpectedly affected messages that were reported with some PAB versions.\n\nOnce we reverted the change and updated the ingester service, PhishER was able to return to normal performance by 20:35 \\(UTC\\).\n\nTo prevent this type of issue in the future, we reviewed the ingester service processes and implemented additional testing for messages reported with the PAB.\u00a0\n\nNo data loss occurred as a result of this issue.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-03-10T20:32:03.311Z",
"resolved_inferred": false,
"started_at": "2026-03-10T15:48:53.428Z",
"state": "postmortem",
"title": "PhishER Showing Incorrect Subject Line for Inbox Messages",
"updated_at": "2026-05-13T17:54:18.473Z",
"url": "https://stspg.io/l0l53rp5677q"
},
{
"body": "From Monday, March 9, 2026, at approximately 19:00 \\(UTC\\) until Tuesday, March 10, 2026, at approximately 17:00 \\(UTC\\), a subset of Workspace customers experienced Secure Webform submission failures.\n\nThis incident was caused when a renewed multi-SAN certificate used by our shared authentication infrastructure \\(Core ESI\\) referenced a newer Certificate Authority \\(CA\\) root. This update caused dependent infrastructure systems running extended-support operating systems to reject authentication requests due to missing trust anchors, leading to TLS validation failures.\n\nTo resolve this issue, the required CA root was deployed to the trust stores of the affected systems and service containers, restoring normal authentication behavior. Failed submissions were successfully replayed. Workspace returned to normal performance by approximately 17:00 \\(UTC\\) on March 10, 2026.\n\nTo prevent this type of issue in the future, we have reviewed the validation and monitoring processes associated with certificate changes, and we are expanding procedures to include checks from dependent systems and re-architecting systems relying on extended support to ensure required CA roots are available.\n\nNo data loss occurred as a result of this issue.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-03-11T10:57:05.473Z",
"resolved_inferred": false,
"started_at": "2026-03-10T00:09:12.539Z",
"state": "postmortem",
"title": "Webform submission failures (UK & EMEA)",
"updated_at": "2026-05-27T21:16:20.187Z",
"url": "https://stspg.io/prh9qqql1gsj"
},
{
"body": "On Thursday, March 5, 2026, from approximately 02:20 until 03:10 \\(UTC\\), some customers experienced expired SSL certificate warnings when accessing their Defend and Prevent admin consoles.\n\nThis incident was caused by an expired certificate for an internal service. To resolve this issue, we deployed a new certificate and updated all affected instances. The Defend and Prevent consoles returned to normal performance by 03:10 \\(UTC\\).\n\nTo prevent this type of issue in the future, we are improving our certificate monitoring.\n\nNo data loss occurred as a result of this issue.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-03-05T12:44:16.511Z",
"resolved_inferred": false,
"started_at": "2026-03-05T02:36:01.441Z",
"state": "postmortem",
"title": "Defend and Prevent Admin Console Certificate Issue",
"updated_at": "2026-03-16T15:50:51.499Z",
"url": "https://stspg.io/rbg9rfh3g2mz"
},
{
"body": "On Tuesday, February 10, 2026, from approximately 13:57 to 16:01 \\(UTC\\), some users experienced delays with the integration service on the EU instance of the PhishER console.\n\nThis incident was caused by a system upgrade that unexpectedly affected the integration service, causing errors in the production environment.\n\nTo resolve this issue, we reverted the update, and the PhishER console returned to normal performance by 16:01 \\(UTC\\).\n\nNo data loss occurred as a result of this issue.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-02-10T15:52:33.108Z",
"resolved_inferred": false,
"started_at": "2026-02-10T14:05:58.464Z",
"state": "postmortem",
"title": "EU PhishER Integration Delay",
"updated_at": "2026-02-13T22:46:46.663Z",
"url": "https://stspg.io/59gg5nmv3nsv"
},
{
"body": "On Tuesday, February 10, 2026, customers in the UK region began experiencing upload delays with Workspace at approximately 11:01 \\(UTC\\), due to instability within a messaging node. The node was isolated at 12:20 \\(UTC\\) to protect cluster integrity and was successfully reintegrated at 13:40 \\(UTC\\), restoring normal upload processing shortly after.\n\nBeginning at approximately 12:10 \\(UTC\\), a subset of Workspace instances across multiple regions became slow or temporarily unresponsive due to memory exhaustion on their underlying virtual machines. As each Workspace instance operates independently, the impact varied by customer.\u00a0\n\nTo resolve this issue, we identified and progressively redeployed the affected instances, starting at 14:37 \\(UTC\\), with the majority restored shortly after. The final affected instance was confirmed healthy at 19:27 \\(UTC\\).\n\nOur analysis confirmed that abnormal memory consumption by the Azure Monitoring Agent caused the host memory exhaustion observed on impacted instances. An infrastructure update scan occurred earlier that morning, at 08:22 \\(UTC\\). Current evidence indicates that this scan triggered an underlying defect in the monitoring agent. The scan itself was non-disruptive in design and was a platform security procedure, not intended as a customer-impacting maintenance activity.\n\nWe confirmed that all systems were stable on February 10 and completed a deep review of data validation on February 11, 2026. No data loss or security compromise occurred.\n\nTo prevent this type of impact in the future, we are strengthening monitoring safeguards around host memory behaviour, enhancing messaging cluster resilience, reinforcing update governance controls, and formalising rapid restoration procedures.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-02-11T14:08:35.581Z",
"resolved_inferred": false,
"started_at": "2026-02-10T12:25:54.000Z",
"state": "postmortem",
"title": "Workspace degraded performance",
"updated_at": "2026-02-25T15:38:39.144Z",
"url": "https://stspg.io/s17bl5d9gz4v"
},
{
"body": "On Wednesday, February 4, 2026, from approximately 15:30 to 16:00 \\(UTC\\), some US users experienced email delivery delays in the Defend product.\u00a0\n\nThis issue was caused by a brief networking fluctuation, resulting in email delivery delays of up to 10 minutes during the affected period. Once the cloud environment returned to healthy operating levels, Defend returned to normal performance, and all delayed emails were delivered by 16:00 \\(UTC\\).\n\nNo data loss occurred as a result of this issue.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-02-04T23:00:15.961Z",
"resolved_inferred": false,
"started_at": "2026-02-04T16:17:27.479Z",
"state": "postmortem",
"title": "Delayed inbound email delivery - US Only",
"updated_at": "2026-02-06T16:29:12.946Z",
"url": "https://stspg.io/5jv6hdbggqqn"
},
{
"body": "On Tuesday, January 27, 2026, from approximately 14:46 \\(UTC\\) to 16:22 \\(UTC\\), some users experienced issues with loading the **Library** tab of the Learner Experience \\(LX\\).\u00a0\n\nThis incident was caused when an update to the **Library** tab contained a code change that conflicted with another recent update. The conflicting code changes unexpectedly caused an error on the LX that was not detected by our automated testing and monitoring until both updates were deployed to the production environment. To resolve this issue, we rolled back the updates, and the LX returned to normal performance by 16:22 \\(UTC\\).\n\nTo prevent this type of issue in the future, we updated the code to prevent the error when the update is redeployed. We also improved the testing and monitoring protocols.\u00a0\n\nNo data loss occurred as a result of this issue.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-01-27T18:04:50.277Z",
"resolved_inferred": false,
"started_at": "2026-01-27T16:20:51.155Z",
"state": "postmortem",
"title": "KSAT - Library Tab not Loading in Learner Experience",
"updated_at": "2026-01-30T15:11:20.063Z",
"url": "https://stspg.io/m2p2yjvn01nt"
},
{
"body": "On January 22, 2026, from approximately 19:30 \\(UTC\\) to January 23, 2026, at approximately 05:00 \\(UTC\\), a significant service disruption in Microsoft\u2019s North American infrastructure triggered a series of service degradations across KnowBe4 products. This Microsoft service disruption primarily affected Exchange Online connectivity and the responsiveness of the Microsoft Graph API and the Office JS API.\n\nThe incident created a two-fold impact on our environment. First, Microsoft\u2019s SMTP relays began rejecting connections, which prevented our secure gateways from returning processed mail to customer environments. As a result, Defend and Protect began safely queuing messages and reached approximately 530,000 emails held in a secure retry state. Prevent also had partial disruptions in which moderation messages could not be transmitted. These disruptions led to delivery delays and triggered preconfigured default actions.\n\nWhile the SMTP relays were down, API instability also affected our suite of Microsoft Outlook add-ins. Both the Phish Alert Button \\(PAB\\) and the KnowBe4 Email Security add-in experienced functional failures and extreme latency because they could not retrieve necessary data from Microsoft\u2019s backend to process reports and display security nudges. At the same time, the deployment center recorded failed installations for new customers, since the system-generated \"handshake\" that is required to verify transport rules was blocked by the same Microsoft relay failures.\n\nAs Microsoft began restoring its North American services on January 23, 2026, at around 01:40 \\(UTC\\), our systems initiated automated recovery protocols. Using exponential backoff logic, our gateways began automatically clearing the mail queues. By 05:00 \\(UTC\\), the Defend and Protect services returned to full performance, queues were fully processed, and all deployment health checks for Defend and Prevent were confirmed successful.\n\nNo data loss occurred as a result of this issue.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-01-23T12:27:41.273Z",
"resolved_inferred": false,
"started_at": "2026-01-22T20:15:41.550Z",
"state": "postmortem",
"title": "Multiple outages due to Microsoft 365 incident",
"updated_at": "2026-02-02T17:36:54.736Z",
"url": "https://stspg.io/lhc78b35y3sn"
},
{
"body": "On Wednesday, January 21, 2026, from approximately 14:59 to Thursday, January 22, 2026, at approximately 10:11 \\(UTC\\), some customers experienced delays with outbound email delivery for UK instances of Prevent, Defend, and Protect.\n\nThis incident was caused by an external email reputation service provider that temporarily restricted one of our outbound email IP addresses. Some messages routed through this restricted address experienced temporary delays while we attempted delivery using other sending addresses. To resolve this issue, our team reached out to the service provider to have the IP address taken off their blocklist. Once the restriction was lifted at 17:33 \\(UTC\\), our systems automatically retried the remaining queued messages, and outbound email delivery returned to normal performance by January 22, 2026, at 10:11 \\(UTC\\).\n\nTo prevent this type of issue in the future, we are investigating additional mitigations for external reputation factors. We are also enhancing our existing monitoring to identify and respond to outbound email reputation changes more quickly.\n\nNo data loss occurred as a result of this issue.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-01-22T12:05:40.359Z",
"resolved_inferred": false,
"started_at": "2026-01-21T18:18:26.673Z",
"state": "postmortem",
"title": "Delayed Email Delivery (UK)",
"updated_at": "2026-02-04T20:49:32.849Z",
"url": "https://stspg.io/tmkgc9tkpzkj"
},
{
"body": "This incident has been resolved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-12-29T12:10:58.242Z",
"resolved_inferred": false,
"started_at": "2025-12-29T10:37:22.504Z",
"state": "resolved",
"title": "Delayed Webform Submissions",
"updated_at": "2025-12-29T12:10:58.259Z",
"url": "https://stspg.io/9cp5j3j8lxq0"
},
{
"body": "On Tuesday, December 23, 2026, from approximately 18:54 to 19:33 \\(UTC\\), some customers experienced delays with their web form submissions.\n\nThis issue was caused by unexpected processing delays in the network nodes that process web forms. To resolve this issue, we redeployed the infrastructure, and web form submissions returned to normal performance by 19:33 \\(UTC\\).\n\nTo prevent this type of issue in the future, we are enhancing web form monitoring and implementing automated corrections in the event a similar error occurs again.\n\nNo data loss occurred as a result of this issue.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-12-24T19:54:01.090Z",
"resolved_inferred": false,
"started_at": "2025-12-23T19:09:20.637Z",
"state": "postmortem",
"title": "Delayed Webform Submissions",
"updated_at": "2026-01-21T16:57:13.566Z",
"url": "https://stspg.io/hfgg6vlfqlvt"
},
{
"body": "On Monday, December 22, 2025, from approximately 16:12 to 16:50\u00a0 \\(UTC\\), some customers experienced errors when attempting to log in and log out across the KSAT console.\u00a0\n\nThis issue was caused by a cryptography library update required for a new feature. To resolve this issue, we reverted the update, and the KSAT console returned to normal performance by 16:50 \\(UTC\\).\n\nTo prevent this type of issue in the future, we have updated our deployment workflows to require sanitation tests before new deployments are sent.\n\nNo data loss occurred as a result of this issue.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "critical",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-12-22T17:04:48.359Z",
"resolved_inferred": false,
"started_at": "2025-12-22T16:20:38.838Z",
"state": "postmortem",
"title": "KSAT - 500 error when logging in",
"updated_at": "2026-01-20T17:25:04.932Z",
"url": "https://stspg.io/q6gstbqfbd5s"
},
{
"body": "On Monday, December 15, 2025, from approximately 18:39 to 19:32 \\(UTC\\), customers were unable to load user data across all regions of the KSAT console.\u00a0\n\nThis issue was immediately identified by our monitoring systems, and the root cause was traced to a database migration during a product deployment. During the migration, a column was intentionally removed from the database as part of a feature deprecation task. However, due to a caching inconsistency between the application images and the database structure, a synchronization gap was created.\n\nOur schema cache system retained old references to the deleted column in the running application images. Because the old application image was still querying for the deleted column and the database no longer contained it, active instances began throwing errors. This issue resulted in a failure to load user data across all regions.\n\nOur engineering team restored service by re-adding the deleted column to both tables. This action resolved the database query inconsistency and allowed the system to load user data correctly. The KSAT console returned to normal performance by 19:32 \\(UTC\\).\n\nNo data loss occurred as a result of this issue.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-12-16T13:59:41.568Z",
"resolved_inferred": false,
"started_at": "2025-12-15T18:54:13.606Z",
"state": "postmortem",
"title": "Users Lists Not Populating",
"updated_at": "2026-01-20T19:10:07.902Z",
"url": "https://stspg.io/lg0zr69y6j7d"
},
{
"body": "The Azure Front Door issue is no longer impacting services. This incident is now resolved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-11T12:29:42Z",
"resolved_at": "2025-12-15T12:32:40.485Z",
"resolved_inferred": false,
"started_at": "2025-12-15T09:49:24.256Z",
"state": "resolved",
"title": "Multiple Products Intermittently Unavailable Due to Azure Front Door Outage (UK Only)",
"updated_at": "2025-12-15T12:32:40.504Z",
"url": "https://stspg.io/c03cgmlsyjvn"
},
{
"body": "This incident has been resolved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-04T10:17:20Z",
"resolved_at": "2025-12-11T20:02:12.009Z",
"resolved_inferred": false,
"started_at": "2025-12-11T12:54:58.918Z",
"state": "resolved",
"title": "Multiple Products Intermittently Unavailable Due to Azure Front Door Outage (UK Only)",
"updated_at": "2025-12-11T20:02:12.024Z",
"url": "https://stspg.io/qsrf7789dmk5"
}
]
}