{
"vendor": "Auvik",
"slug": "auvik",
"platform": "statuspage",
"status_url": "https://status.auvik.com",
"last_checked": "2026-09-16T12:28:20Z",
"last_state": "ok",
"history_backfilled": true,
"first_watched": "2026-09-04T07:06:16Z",
"incidents": [
{
"body": "The incident has been fully resolved, and all services are operating normally.\n\nCustomers should no longer experience any related issues. If you continue to experience problems, please don't hesitate to contact Auvik Support.\n\nWe will provide a Root Cause Analysis (RCA) once it is available.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "critical",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-09-03T08:18:52.526-04:00",
"resolved_inferred": false,
"started_at": "2026-09-03T07:12:53.072-04:00",
"state": "resolved",
"title": "Incident Title",
"updated_at": "2026-09-03T08:18:52.542-04:00",
"url": "https://stspg.io/sk7gl55zfj27"
},
{
"body": "# Service Disruption - Availability Issues on the AU1 Cluster\n\n## Root Cause Analysis\n\n### Duration of the incident\n\nDiscovered: Aug 27, 2026  12:20 - UTC  \nResolved:     Aug 27, 2026  13:15 - UTC\n\n### Customer impact\n\nClients hosted in the AU1 region experienced a temporary interruption of Auvik services. Depending on the client, monitoring, alerting, collector connectivity, and access to the web interface were unavailable for up to approximately one hour.\n\n### Cause\n\nA cascading failure in core platform services caused multiple service components to restart, including a central coordination component. This resulted in widespread client-environment restarts across the AU1 region.\n\n### Effect\n\nAffected client environments became temporarily unavailable while services restarted. The platform recovered automatically. Post-recovery verification confirmed normal service operation.\n\n### Future consideration\\(s\\)\n\n* Continue evaluating resiliency improvements for core service coordination and recovery behavior following large-scale restarts.\n* Continue using post-recovery validation checks to confirm client services return to expected operation after a broad service restart.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-08-27T09:30:25.575-04:00",
"resolved_inferred": false,
"started_at": "2026-08-27T08:46:10.771-04:00",
"state": "postmortem",
"title": "Service Disruption -  AU1 Cluster Sites",
"updated_at": "2026-08-28T11:24:11.562-04:00",
"url": "https://stspg.io/8ztjw0r1mtth"
},
{
"body": "# Service Degradation - Legacy Alerts Notification\n\n## Root Cause Analysis\n\n### Duration of the incident\n\nDiscovered: Aug 17, 2026  13:15 - UTC  \nResolved:     Aug 21, 2026  14:12 - UTC\n\n### Customer impact\n\nSome customers using Legacy Alerts experienced inaccurate alert notifications, including false-positive alerts and unintended email notifications. During remediation, some legacy alert processing was delayed, and alert visibility in the user interface temporarily lagged. V2 Alerts remained operational throughout the incident.\n\n### Cause\n\nA product release intended to retire older default Legacy Alert configurations for new environments also removed supporting alert conditions that were still required by some existing customers with customized Legacy Alerts. Without the complete conditions, certain Legacy Alerts could trigger inaccurately. The remediation required reprocessing affected alerts, which temporarily increased processing load and contributed to delayed alert visibility, alert flapping, and excessive notifications.\n\n### Effect\n\nAffected Legacy Alerts could trigger when the underlying condition was not present, re-trigger after being paused, or send notifications more broadly than intended. The resulting in-service lag also temporarily delayed displaying some alerts. Engineering restored the required alert configuration, removed the unintended broad email association, increased processing capacity, and proactively corrected additional affected customer environments.\n\n### Future consideration\\(s\\)\n\n\u2022 Add monitoring for abnormal volumes of Legacy Alerts so unexpected alert activity can be detected earlier.\n\n\u2022 Expand pre-release and end-to-end testing for existing customized Legacy Alerts, including alert triggering, notification delivery, and post-change recovery behavior.\n\n\u2022 Strengthen canary-release review so alerting anomalies are investigated and understood before wider deployment.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-08-20T10:11:24.839-04:00",
"resolved_inferred": false,
"started_at": "2026-08-17T13:53:50.634-04:00",
"state": "postmortem",
"title": "Inaccurate notifications being sent for legacy alerts",
"updated_at": "2026-08-24T13:57:17.943-04:00",
"url": "https://stspg.io/701sb7knmbl9"
},
{
"body": "The incident has been fully resolved, and all services are operating normally.\n\nCustomers should no longer experience any related issues. If you continue to experience problems, please don't hesitate to contact Auvik Support.\n\nWe will provide a Root Cause Analysis (RCA) once it is available.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-08-17T12:42:43.026-04:00",
"resolved_inferred": false,
"started_at": "2026-08-17T11:00:40.735-04:00",
"state": "resolved",
"title": "Issue with Legacy Alert Delivery",
"updated_at": "2026-08-17T12:42:43.044-04:00",
"url": "https://stspg.io/lcbjryjg4jzj"
},
{
"body": "# Service Disruption - SNMPv3 Monitoring Loss Following Collector Upgrade\n\n## Root Cause Analysis\n\n### Duration of the incident\n\nDiscovered: Apr 13, 2026 20:00 - UTC  \nResolved:     Apr 13, 2026 22:32 - UTC\n\n### Customer impact\n\nCustomers experienced a loss of monitoring data from only devices using specific SNMPv3 configurations. While devices remained online and reachable, monitoring data was not collected, resulting in reduced visibility across environments.\n\n### Cause\n\nA recent collector upgrade introduced changes to encryption handling that affected support for certain legacy SNMPv3 configurations. This resulted in failures when attempting to collect data from devices configured that way.\n\n### Effect\n\nMonitoring data collection failed for affected devices across multiple clusters. This led to a noticeable drop in available device metrics and visibility, despite no loss of connectivity to the devices themselves.\n\n### Future consideation\\(s\\)\n\n* Expand test coverage to include a broader range of SNMP configurations\n* Improve monitoring to detect drops in data collection more proactively\n* Strengthen validation processes for major upgrades and dependency changes\n* Implement additional safeguards to identify compatibility issues prior to release",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-04-13T22:33:40.666-04:00",
"resolved_inferred": false,
"started_at": "2026-04-13T20:32:11.502-04:00",
"state": "postmortem",
"title": "Emergency Collector Rollback",
"updated_at": "2026-04-16T18:14:54.952-04:00",
"url": "https://stspg.io/16gkwgd9vr90"
},
{
"body": "# Service Degraded - Some Customers Experiencing Login Issues After Maintenance\n\n## Root Cause Analysis\n\n### Duration of the incident\n\nDiscovered: Mar 7, 2026 18:42 - UTC  \nResolved:     Mar 9, 2026 21:10 - UTC\n\n### Cause\n\nFollowing a platform upgrade, an internal reconciliation process incorrectly identified a subset of user records as deleted. This occurred due to an error condition in a synchronization component that failed to process all expected user data during reconciliation.\n\n### Effect\n\nAffected users were able to log in to the platform but were unable to access their expected sites or tenants due to missing user authorizations. In some cases, API access using previously generated API keys was also affected.\n\nMonitoring, alerting, and data collection services continued to function normally throughout the incident.\n\n### Action taken\n\nEngineering paused the process that spreads the changes across the platform to prevent further impact.\n\nUser access was then restored using backup data taken prior to the upgrade. This allowed the team to recover user access, permissions, and related access keys.\n\nAfter restoration was completed and access was verified across all environments, the incident was declared resolved.\n\n### Future consideration\\(s\\)\n\n* Implement additional monitoring to detect unexpected changes in user or authorization records.\n* Add safeguards within reconciliation processes to prevent large-scale unintended deletions.\n* Improve post-upgrade validation checks to identify abnormal record changes earlier.\n* Enhance disaster recovery tooling and backup retrieval processes to speed up restoration efforts.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-03-09T17:26:06.121-04:00",
"resolved_inferred": false,
"started_at": "2026-03-08T15:10:43.417-04:00",
"state": "postmortem",
"title": "Service Degradation \u2013 Some Customers Experiencing Login Issues After Maintenance",
"updated_at": "2026-03-17T10:52:08.942-04:00",
"url": "https://stspg.io/y5q08g56cky8"
},
{
"body": "# Service Disruption - The EU2 cluster was unavailable \n\n## Root Cause Analysis\n\n### Duration of the incident\n\nDiscovered: Dec 15, 2025 20:00 \u2013 UTC  \nResolved:     Dec 16, 2025 02:00 \u2013 UTC\n\n### Customer impact\n\nDuring the incident window, customers hosted in the EU2 region experienced intermittent service degradation. This included slower system responsiveness, temporary inconsistencies in monitoring data, and brief periods where alerts may have been delayed or inaccurate.  \nMost customers regained access as services were progressively restored, and complete stability was confirmed before the incident was closed.\n\n### Cause\n\nThe incident was caused by an elevated load in the EU2 service environment, resulting in an uneven workload distribution across backend resources. As the load increased, automated recovery mechanisms were unable to stabilize the environment fully, necessitating a controlled restart of the regional service to restore normal operations.\n\n### Effect\n\nThe imbalance led to reduced service performance and temporary unavailability for some customers until recovery actions were completed. Engineering teams were required to intervene to safely restart the affected region and validate service health before returning operations to normal.\n\n### Future consideation\\(s\\)\n\n* Improve automated workload balancing to absorb regional load increases.\n* Strengthen early indicators for backend saturation to enable earlier intervention.\n* Refine operational procedures to further reduce recovery time in similar scenarios.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-12-15T20:57:36.921-05:00",
"resolved_inferred": false,
"started_at": "2025-12-15T18:28:18.629-05:00",
"state": "postmortem",
"title": "Emergency maintenance for EU2",
"updated_at": "2025-12-22T03:52:15.120-05:00",
"url": "https://stspg.io/jgsqkrqz30t6"
},
{
"body": "# Service Degraded - Duplicated Devices for Meraki Switch Stacks\n\n## Root Cause Analysis\n\n### Duration of the incident\n\nDiscovered: Dec 15, 2025 08:34 \u2013 UTC  \nResolved:     Dec 15, 2025 13:20 \u2013 UTC\n\n### Customer impact\n\nCustomers with Meraki switch stacks across all clusters experienced:  \nDuplicate Meraki switches are appearing in the device inventory, including entries without management IP addresses.  \nInability to manually delete duplicates, as they were automatically recreated.  \nConfusing or inaccurate device and stack representations.  \nTemporary inflation of billable device counts for some tenants \\(billing data captured and under review\\).  \nNo other device types were affected.\n\n### Cause\n\nA configuration update intended to improve how Meraki switch stacks are represented in the platform caused unintended behavior.  \nFor tenants already using Meraki stacking, the system was unable to match newly processed stack data to existing switches consistently. This caused already-stacked switches to be incorrectly interpreted as new devices, generating duplicate inventory entries. Because the underlying synchronization logic reinforced these entries, any customer deletion of duplicates did not persist, and duplicates reappeared.\n\n### Effect\n\nThe issue led to inaccurate device inventories and misleading switch-stack views for customers using Meraki stacking, causing confusion in device management and increasing support tickets. Internal teams experienced additional operational overhead as they investigated the duplication patterns and validated safe removal procedures. Complete remediation required coordinated cleanup across all clusters to restore accurate device records and ensure no further duplicates were generated.\n\n### Future consideration\\(s\\)\n\n* Strengthen validation of device and stack metadata before inventory updates.\n* Improve testing coverage using a production-like stacked device configuration.\n* Add monitoring and alerting for abnormal changes in device counts.\n* Use safer rollout controls and feature flags for changes affecting device inventory logic.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-12-15T13:23:04.412-05:00",
"resolved_inferred": false,
"started_at": "2025-12-15T09:35:29.220-05:00",
"state": "postmortem",
"title": "Meraki Switch stacks creating duplicate devices in the product",
"updated_at": "2025-12-19T11:28:29.934-05:00",
"url": "https://stspg.io/t6wwffd07tjk"
},
{
"body": "# Service Degraded - FortiGate firewalls duplicated in the EU2 cluster, causing a flood of alerts.\n\n## Root Cause Analysis\n\n### Duration of the incident\n\nDiscovered: Dec 11, 2025 14:00 \u2013 UTC  \nResolved:     Dec 17, 2025 06:45 \u2013 UTC\n\n### Customer impact\n\nClients with FortiGate firewalls in the EU2 region experienced:\n\n* A significant increase in \u201cdevice offline\u201d alerts due to status flapping between duplicate and original device entries.\n* Duplicate firewalls appear in inventory views, leading to confusion in device management.\n* Inaccurate device statistics and monitoring gaps.\n* In some cases, customers attempted to delete duplicates, which resulted in the loss of historical data on the original device.\n\nNo other device types or regions were affected.\n\n### Cause\n\nA system change intended to improve how high-availability firewall configurations were processed introduced unexpected behavior. Under certain conditions, the platform was unable to detect when multiple records referenced the same firewall device consistently. As a result, some firewalls were incorrectly detected as new devices, causing duplicate entries and inconsistent operational status reporting.\n\n### Effect\n\nThe issue resulted in elevated alert volumes across impacted tenants and caused instability in device-status reporting for the affected firewalls. Customers experienced increased operational overhead as they reviewed unexpected alerts and duplicate device entries, and engineering teams were required to intervene to restore accurate device associations and remove the duplicated records.fFuture consideration\\(s\\)\n\n* Strengthen validation of device-supplied information before it affects inventory or monitoring.\n* Implement alerting for sudden spikes in offline/online transitions to detect similar issues earlier.\n* Improve observability of device lifecycle behavior to identify anomalies proactively.\n* Use feature-flagged or staged rollouts for changes that affect device processing logic.\n* Incorporate production-like data into pre-deployment testing to better anticipate unexpected device behaviors.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-12-11T18:46:29.027-05:00",
"resolved_inferred": false,
"started_at": "2025-12-11T09:30:41.924-05:00",
"state": "postmortem",
"title": "Duplicate FortiNet firewalls created on the EU2 cluster.",
"updated_at": "2025-12-19T11:19:26.806-05:00",
"url": "https://stspg.io/yq5p5r9dms36"
},
{
"body": "The incident has been fully resolved, and all services are operating normally.\n\nCustomers should no longer experience any related issues. If you continue to experience problems, please don't hesitate to contact Auvik Support.\n\nWe appreciate your understanding and patience.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-12-05T08:32:57.269-05:00",
"resolved_inferred": false,
"started_at": "2025-12-05T07:50:30.000-05:00",
"state": "resolved",
"title": "Auvik Live Chat is Inaccessible",
"updated_at": "2025-12-05T08:32:57.285-05:00",
"url": "https://stspg.io/1gxz2qhl20kj"
},
{
"body": "Latest update from Okta\nIncident Resolved\n\nDecember 3, 2025 at 8:58am PST\n\nAn issue impacting the core authentication service for customers in Okta Cell US7 has been resolved. Additional root cause information will be available within five business days.\n\nWe will close the status page for Auvik's sites.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-12-03T12:28:00.389-05:00",
"resolved_inferred": false,
"started_at": "2025-12-03T09:21:09.865-05:00",
"state": "resolved",
"title": "Service Disruption -  Login issues",
"updated_at": "2025-12-03T12:28:00.407-05:00",
"url": "https://stspg.io/19npfpys2xh2"
},
{
"body": "# Service Degraded - Data Not Loading Across All Clusters\n\n## Root Cause Analysis\n\n### Duration of the incident\n\nDiscovered: Nov 21, 2025 17:43 \u2013 UTC  \nResolved:     Nov 21, 2025 19:18 \u2013 UTC\n\n### Customer impact\n\nCustomers across all regions were able to log in, but many were unable to view key data within the platform. This included features such as:\n\n* Network maps\n* Inventory and site lists\n* Dashboards\n* Other views are dependent on permissions and hierarchical data.\n\nThis resulted in degraded usability and limited visibility into managed environments.\n\n### Cause\n\nDuring a routine maintenance task intended to clean up older data records, an incorrect piece of information was unintentionally added to the system. This faulty data prevented a core component\u2014responsible for organizing how customer information is displayed in the platform\u2014from functioning correctly.\n\nBecause this component could not run as expected, many areas of the product that rely on structured data \\(such as maps, dashboards, and inventory views\\) were unable to load. This led to widespread service degradation across all clusters.\n\n### Effect\n\nBecause the system could not correctly load the information needed to display customer environments, several parts of the platform were unable to show data. As a result, many users could log in but were unable to see maps, inventory details, site lists, dashboards, or other information they usually rely on.  \nThis created a degraded experience across all regions, limiting visibility and making it difficult for customers to perform routine monitoring and management tasks until the issue was resolved.\n\n### Future consideation\\(s\\)\n\n* Introduce stronger validation and safeguards to prevent malformed data from being accepted into critical systems.\n* Improve the resilience of backend services so they fail gracefully rather than entering repeated restart cycles.\n* Implement additional monitoring to provide earlier detection of issues affecting data-loading functions.\n* Enhance internal tools used during maintenance and cleanup activities to reduce the risk of unintentional data corruption.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-11-21T14:26:06.441-05:00",
"resolved_inferred": false,
"started_at": "2025-11-21T13:29:10.342-05:00",
"state": "postmortem",
"title": "Service Disruption -  Reduced functionality across all clusters",
"updated_at": "2025-12-04T11:39:01.078-05:00",
"url": "https://stspg.io/9phgpqx8cp1d"
},
{
"body": "The incident has been fully resolved, and all services are operating normally.\n\nCustomers should no longer experience any related issues. If you continue to experience problems, please don't hesitate to contact Auvik Support.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-10-20T06:57:38.371-04:00",
"resolved_inferred": false,
"started_at": "2025-10-20T05:51:11.783-04:00",
"state": "resolved",
"title": "Partial service degradation due to AWS outage",
"updated_at": "2025-10-20T06:57:38.389-04:00",
"url": "https://stspg.io/h5xgd9bd0q4y"
},
{
"body": "# Service Degraded - Sites have lost settings after maintenance\n\n## Root Cause Analysis\n\n### Duration of the incident\n\nDiscovered: Oct 11, 2025 \u2013 12:00 UTC  \nResolved:     Oct 15, 2025 \u2013 22:00 UTC\n\n### Cause\n\nDuring a scheduled system update, a background maintenance process unintentionally removed reference files used to identify stored site configurations. When affected systems restarted after the update, they were unable to locate those configuration files and temporarily appeared as new, empty sites.\n\nThis occurred because the maintenance process was using outdated information when determining which data to clean up safely. The underlying data remained securely stored, but the missing reference files prevented normal access until they were restored.\n\n### Effect\n\nA subset of tenants across multiple clusters temporarily lost access to their site configurations and appeared as newly created environments.  \nCustomers observed missing data and configurations, including previously defined network settings and device details.\n\n### Action taken\n\n_All times are in UTC_\n\n**10/11/2025**\n\n**12:00** \u2013 A scheduled system upgrade began across all clusters.\n\n**15:25** \u2013 The support team received reports from customers that some sites appeared empty or missing data.\n\n**16:00** \u2013 Engineering immediately began investigating and determined this was not related to normal data processing delays.\n\n**10/12/2025**\n\nAdditional reports confirmed that several sites were missing configuration information.  The engineering team confirmed that the original data was still securely stored, but was not being correctly loaded by the system.\n\n**10/13/2025**\n\nThe issue was traced to missing metadata files that help identify stored configurations.  The engineering team began restoring affected sites using the most recent valid configuration data.\n\nAutomated recovery tools were developed to safely restore additional sites and ensure consistent recovery across all clusters.\n\n**10/14/2025**\n\nEngineering verified that configuration data was fully restored and synchronized across supporting services.\n\nThe recovery process was extended to all remaining sites, with validation steps confirming successful restoration.\n\n**10/15/2025**\n\nFinal recovery efforts for the remaining affected clusters were completed..\n\n**22:00** \u2013 All affected sites were confirmed operational with their configurations restored and verified.\n\n### Future consideration\\(s\\)\n\n* Remove dependency of cleanup processes on outdated cluster data sources.\n* Validate all automated cleanup jobs to ensure they do not operate on production clusters.\n* Implement monitoring for missing or corrupted metadata files before deployments.\n* Enhance post-deployment validation to verify the integrity of configuration data across all clusters.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-10-15T13:33:51.847-04:00",
"resolved_inferred": false,
"started_at": "2025-10-15T11:22:07.158-04:00",
"state": "postmortem",
"title": "Clients on US4 cluster are experiencing 500 Errors",
"updated_at": "2025-10-31T10:31:08.394-04:00",
"url": "https://stspg.io/x1w54rqkp966"
},
{
"body": "# Service Disruption - Platform Availability and Login Access \n\n## Root Cause Analysis\n\n### Duration of the incident\n\nDiscovered: Sep 23, 2025 \u2013 17:45- UTC  \nResolved:     Sep 26, 2025 \u2013 13:58 - UTC\n\n### Customer impact\n\nTenants on the CA1 cluster lost their settings, and Auvik was inaccessible for a time.\n\nIntermittent platform errors and degraded performance for some customers.\n\nA temporary issue prevented certain users from logging in via Okta.\n\nDelays in data synchronization in the EU1 region necessitate restarting the cluster.\n\n### Cause\n\nA configuration update intended for specific tenants was applied globally due to a missing query constraint.\n\nThis caused the deletion of tenant and user records in one cluster, leading to cascading synchronization and workload impacts across other clusters.\n\nAs a result, user data was propagated as deletions to the identity provider, temporarily removing affected user accounts.\n\n### Effect\n\nThe CA1 cluster required restoration from a backup, resulting in the temporary unavailability of some tenant data.\n\n43 tenants and 22 user accounts required manual recreation and verification.\n\nAuthentication failures occurred for affected users due to the removal of identity records.\n\nThe EU1 cluster experienced performance degradation as it processed a large data synchronization backlog.\n\nIncreased load briefly impacted the responsiveness of other clusters\u2019 regional services.\n\n### Future consideration\\(s\\)\n\n* Implement a dry-run and confirmation step in internal tooling to validate production commands before execution.\n* Reinforce the change-management process for improved peer visibility and validation.\n* Expand resource and capacity monitoring to identify anomalies earlier and respond proactively.\n* Revisit release control lifecycle practices to remove obsolete configurations after rollout completion.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-09-26T09:58:35.912-04:00",
"resolved_inferred": false,
"started_at": "2025-09-23T13:45:56.147-04:00",
"state": "postmortem",
"title": "Site performance and access issues",
"updated_at": "2025-10-15T13:51:47.284-04:00",
"url": "https://stspg.io/083y9jrc99c4"
},
{
"body": "# Service Degraded - Clients experienced login issues to their site because the URL redirect was not working.\n\n## Root Cause Analysis\n\n  \nA recent update led to more traffic than expected, resulting in simultaneous overload of the same systems. This overloaded them, leading to delays and occasional failures when customers tried to log in. Some customers also experienced issues when trying to start new trials.\n\n### Duration of the incident\n\nDiscovered: Sep 17, 2025 13:50 - UTC  \nResolved:     Sep 18, 2025 21:00 - UTC\n\n### Cause\n\nThe update unintentionally created extra demand on shared systems. As a result, the login process and new trial creation sometimes failed or responded slowly.\n\n### Effect\n\n* Some customers could not log in after entering their password or completing MFA.\n* Occasional slow responses and error messages \\(404/500/502/504\\).\n* New trial sign-ups sometimes failed or were delayed.\n* Internal tools that rely on the same login process also saw intermittent issues.\n\n### Action taken  \n\n* Adjusted system settings to reduce pressure on overloaded services.\n* Closely monitored traffic while making changes to keep the service stable.\n* Applied a temporary workaround to allow new trials to be created reliably.\n* Released a fix to stabilize the login redirect process.\n* Made further tuning changes to spread out demand and reduce load.\n* Continued monitoring until the login and sign-up flows were confirmed to be stable.\n\n### Future consideration\\(s\\)\n\n* Reduce reliance on a single system for login and trial flows by distributing the workload across multiple systems..\n* Enhance monitoring to identify login errors and trial creation issues more promptly.\n* Add safeguards to prevent overload, including traffic limits and fallback options.\n* Test updates under heavier load conditions to catch these issues earlier.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-09-18T17:00:23.307-04:00",
"resolved_inferred": false,
"started_at": "2025-09-17T09:53:56.227-04:00",
"state": "postmortem",
"title": "Service Disruption - Log in issues to Auvik",
"updated_at": "2025-09-29T09:54:48.363-04:00",
"url": "https://stspg.io/lsrwk842p6dl"
},
{
"body": "# Service Degraded - Login and Collector Installation Issues on US3 Cluster\n\n## Root Cause Analysis\n\n### Duration of the incident\n\nDiscovered: Sep 10, 2025 13:17 UTC  \nResolved:     Sep 12, 2025 15:08 UTC\n\n### Cause\n\nTwo separate but overlapping issues contributed to this incident:\n\n* Collector Installation Failures \u2013 Windows collectors were unable to install due to missing service principal credentials on the backend agent server, which prevented successful API calls for subscription data.\n* Login and Redirect Failures \u2013 Following a restart of the US3 frontend, requests for user and tenant data from secondary clusters intermittently failed. This caused login attempts through Okta to hang and product redirects to fail.\n\n### Effect\n\n* Users attempting to log in via Okta were unable to complete authentication and access tenants.\n* Some sites experienced 500 errors when attempting to access dashboards.\n* Windows collector installations via GUI and CLI failed, preventing the deployment of new collectors.\n\n### Action taken\n\n_All times are in UTC_\n\n  \n**09/10/2025**\n\n**13:17** \u2013 Users report inability to log in to Auvik Production through Okta.\n\n**13:28** \u2013 Errors in US3 frontend logs identified relating to user/tenant data queries.\n\n**13:37** \u2013 Engineering suspends frontend deployment in the US3 cluster.\n\n**13:39** \u2013 Issues confirmed across secondary clusters; logs analyzed for root cause.\n\n**14:54** \u2013 Identified that the frontend redirect service could not fetch required tenant data; feature flag disabled to restore functionality.\n\n**15:01** \u2013 Engineering confirms that the workaround restores login functionality while monitoring tenant recovery.  The incident is resolved on the status page.\n\n**09/12/2025**\n\n**15:08** \u2013 Feature flag re-enabled after system recovery; frontend services reconciled successfully. Incident fully resolved.\n\n### Future consideration\\(s\\)\n\n* Improve validation of service principal credentials to prevent collector installation failures.\n* Enhance monitoring and alerting around login and redirect workflows to detect tenant query failures earlier.\n* Review and refine feature flag rollout procedures to minimize dependency risks and ensure optimal deployment.\n* Continue improving how customer data is distributed across clusters so that queries run more efficiently and reliably, even during high system load or maintenance events.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-09-10T13:26:51.000-04:00",
"resolved_inferred": false,
"started_at": "2025-09-10T12:43:33.677-04:00",
"state": "postmortem",
"title": "Users are having issues connecting to Auvik",
"updated_at": "2025-09-24T20:19:54.565-04:00",
"url": "https://stspg.io/mc885g7yw5lw"
},
{
"body": "The access issues for clients on the US4 cluster have been fully resolved, and services are operating as expected.\n\nImpact: \nCustomers should no longer experience any related issues. If you continue to experience issues, please report them to Auvik Support.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-08-27T19:06:55.799-04:00",
"resolved_inferred": false,
"started_at": "2025-08-27T17:30:17.548-04:00",
"state": "resolved",
"title": "US4 clients are not accessible",
"updated_at": "2025-09-24T22:37:45.409-04:00",
"url": "https://stspg.io/lnwns1f9qndc"
},
{
"body": "# Service Disruption - Intermittent Availability & Performance Issues Across Multiple Clusters\n\n## Root Cause Analysis\n\n### Duration of the incident\n\nDiscovered: Aug 25, 2025 18:00 - UTC  \nResolved:     Aug 29, 2025 14:00 - UTC\n\n### Cause\n\nA configuration rollout unexpectedly generated a large number of configuration entries, which propagated across tenants. This resulted in excessive background processing and memory pressure in core services. The strain led to degraded performance, instability, and in some cases, brief service crashes across clusters.\n\n### Effect\n\nCustomers experienced:\n\n* Intermittent access and sign-in issues in several regions\n* Slow page loads and missing/delayed alert notifications\n* Errors or gaps in specific dashboard and visualization views\n* Temporary unavailability for a small number of tenants\n\n### Action taken\n\n_All times are in UTC_\n\n**08/25/2025**  \n**18:00** \u2014 Rollout halted after error rates increased.  \n**19:00** \u2014 Targeted service restarts restored partial availability.  \n**22:00** \u2014 Added backend capacity and began controlled rollouts.\n\n**08/26\u201308/28/2025**\n\nContinued staged rollouts with adjusted capacity.  \nCleaned up configuration entries for affected tenants.  \nTuned resource allocations for read/permissioning services\n\n**08/29/2025**\n\n**14:00** \u2014 All clusters stabilized; monitoring confirmed normal performance.\n\n### Future consideration\\(s\\)\n\n* Enhance autoscaling and resource thresholds for services under heavy background processing.\n* Add scale-aware pre-deployment validation for configuration rollouts.\n* Refine monitoring to surface customer-visible issues earlier.\n* Expand operational runbooks for rollback and tenant recovery.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-08-27T10:13:51.026-04:00",
"resolved_inferred": false,
"started_at": "2025-08-25T14:32:03.498-04:00",
"state": "postmortem",
"title": "Service Disruption -  Auvik clients are experiencing a disruption of services - Multiple Clusters",
"updated_at": "2025-09-02T12:39:33.579-04:00",
"url": "https://stspg.io/3rk0kw6jl8zw"
},
{
"body": "# Performance Degraded - Site dropdown not working for several clusters\n\n## Root Cause Analysis\n\n### Duration of the incident\n\nDiscovered: Aug 25, 2025 \u2013 15:40 UTC  \nResolved:     Aug 25, 2025 \u2013 16:45 UTC\n\n### Cause\n\nA recent system update introduced a change that depended on data not yet available in our production environment. As a result, the site selector \\(drop-down\\) was unable to load correctly, which prevented users from switching between sites until the update was rolled back.\n\n### Effect\n\nCustomers were unable to switch sites in the platform, which limited access to some account information and historical data.\n\n### Action taken\n\n_All times are in UTC_\n\n**08/25/2025**\n\n**15:40** \u2013 Update deployed across production clusters.\n\n**16:17** \u2013 Customer reports received that the site drop-down was not loading.\n\n**16:26** \u2013 Engineering identified the change that introduced the issue.\n\n16**:27** \u2013 Incident declared; response team assembled.\n\n**16:40** \u2013 Deployment was rolled back.\n\n**16:45** \u2013 Service restored; site drop-down working as expected.  Incident resolved.\n\n###   \nFuture consideration\\(s\\)\n\n* Add extra checks before updates to ensure all required data is available in production.\n* Improve safeguards so that unrelated changes are not deployed together.\n* Build fallback mechanisms to ensure that new features do not impact existing functionality.\n* Strengthen monitoring to detect and alert on similar issues in the future quickly.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-08-25T12:58:53.882-04:00",
"resolved_inferred": false,
"started_at": "2025-08-25T12:58:53.838-04:00",
"state": "postmortem",
"title": "Site Dropdown and permission issues in the Auvik UI",
"updated_at": "2025-09-10T13:59:53.332-04:00",
"url": "https://stspg.io/r1x66knlhw3s"
},
{
"body": "# Service Degraded - Clients on the EU1 cluster using V2 alerting are not reviewing device alerts.\n\n## Root Cause Analysis\n\n### Duration of the incident\n\nDiscovered: Aug 21, 2025 23:47 - UTC  \nResolved:     Aug 22, 2025 12:00 - UTC\n\n### Cause\n\nA change to the alert-processing timing logic introduced a defect where time windows did not close properly. This prevented events from being processed promptly, causing alerts to queue up and delaying their delivery to the user interface.\n\n### Effect\n\nCustomers on the EU1 cluster using V2 alerting experienced delays in reviewing device alerts, with some alerts being delayed by up to 12 hours.\n\n### Action taken\n\n_All times are in UTC_\n\n**08/21/2025**\n\n**23:47** \u2013 Alert processing began lagging; backlog started building.\n\n**08/22/2025**\n\n**06:00** \u2013 Incident declared; engineers engaged to investigate.\n\n**12:00** \u2013 Adjustments made to processing pipeline; backlog cleared; all delayed alerts reprocessed; service restored.\n\n### Future consideration\\(s\\)\n\n* Add monitoring for time-window stalls and backlog growth.\n* Expand testing to cover out-of-order and skewed event scenarios.\n* Strengthen rollback plans for all future alert-processing changes.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-08-22T06:36:49.497-04:00",
"resolved_inferred": false,
"started_at": "2025-08-22T05:32:22.637-04:00",
"state": "postmortem",
"title": "Service Degraded - Clients on the EU1 cluster using V2 alerting are not reviewing device alerts",
"updated_at": "2025-09-08T09:39:01.878-04:00",
"url": "https://stspg.io/kf7r0frdl94l"
},
{
"body": "# Performance Degraded - Sites on the EU2 cluster are slow to load\n\n## Root Cause Analysis\n\n### Duration of the incident\n\nDiscovered: Aug 20, 2025 09:27 \u2013 UTC  \nResolved:     Aug 20, 2025 10:05 \u2013 UTC\n\n### Cause\n\nOne application component handling user requests on EU2 became unhealthy and stopped responding normally.\n\n### Effect\n\nSome EU2 customers experienced slow page loads and intermittent timeouts in the web experience during the incident window.  \nAction taken\n\n_All times are in UTC_\n\n**08/20/2025**\n\n**09:27** \u2013 Potential slowness on EU2 reported; investigation initiated.\n\n**09:35** \u2013 Elevated errors observed; incident declared; response team engaged.\n\n**09:45** \u2013 Application components restarted to restore service.\n\n**10:05** \u2013 Service performance restored.\n\n**10:07** \u2013 Validation confirmed recovery; incident closed.\n\n### Future consideration\\(s\\)\n\n* Complete an investigation into the component crash and address any defects found.\n* Enhance monitoring/alerting for rising request latency and timeout errors to detect earlier.\n* Review deployment and health-check safeguards to auto-recover unresponsive components safely.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-08-20T10:09:21.589-04:00",
"resolved_inferred": false,
"started_at": "2025-08-20T09:50:24.176-04:00",
"state": "postmortem",
"title": "Performance Issue - Slowness on site on the EU2 cluster",
"updated_at": "2025-09-08T09:34:40.149-04:00",
"url": "https://stspg.io/vwdtzhn64n7d"
},
{
"body": "# Performance Degraded - Clients on the US2 cluster are slow to load\n\n## Root Cause Analysis\n\n### Duration of the incident\n\nDiscovered: Aug 19, 2025 \u2013 13:10 UTC  \nResolved:    Aug 19, 2025 \u2013 15:30 UTC\n\n### Cause\n\nRecent configuration changes to backend data replication caused a surge in database writes. This increased CPU utilization across all clusters, but while other clusters recovered, the US2 database instance did not. The elevated CPU load persisted for over 24 hours, which led to customer-facing slowness when loading sites on the US2 cluster.\n\n### Effect\n\nCustomers on the US2 cluster experienced significantly slower site load times in the Auvik UI. This impacted demos, trials, and production users, resulting in degraded user experience until resolution.\n\n### Action taken\n\n_All times are in UTC_\n\n**08/19/2025**\n\n**13:10** \u2013 Sales reported demo site loading issues on US2.\n\n**13:22** \u2013 Engineering identified elevated CPU usage on the US2 database.\n\n**13:27** \u2013 Investigation into DB performance began.\n\n**13:57** \u2013 Confirmed that US2 had remained at 100% CPU since Aug 18.\n\n**14:25** \u2013 Troubleshooting efforts to recover performance begin.\n\n**14:36** \u2013 Proposal made to scale up resources for the DB.\n\n**15:00** \u2013 Decision made to upgrade the US2 database instance type.\n\n**15:09** \u2013 Database instance size increased.\n\n**15:22** \u2013 Read/write latency returned to normal.\n\n**15:30** \u2013 US2 UI performance confirmed as fully recovered.\n\n### Future consideration\\(s\\)\n\n* Enhance database CPU monitoring alerts to ensure visibility into leading indicators.\n* Improve alerting for customer-impacting issues such as UI slowness.\n* Conduct proactive reviews of cluster resource utilization to identify potential bottlenecks.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-08-19T11:30:00.000-04:00",
"resolved_inferred": false,
"started_at": "2025-08-19T11:30:00.000-04:00",
"state": "postmortem",
"title": "Performance Issue - Sites on US2 cluster are slow to load",
"updated_at": "2025-08-25T19:32:40.480-04:00",
"url": "https://stspg.io/gynnswxpd4jv"
}
]
}