{
"vendor": "Dixa",
"slug": "dixa",
"platform": "statuspage",
"status_url": "https://status.dixa.io",
"last_checked": "2026-09-16T12:28:20Z",
"last_state": "ok",
"history_backfilled": true,
"first_watched": "2026-09-04T07:06:16Z",
"incidents": [
{
"body": "### **Post-Mortem: Advanced Insight & Miuros Degraded Performance**\n\n**Date of Incident:** July 28, 2026\n\n**Duration:** 12:47 PM \u2013 1:15 PM CEST \\(28 minutes\\)\n\n**Severity:** Medium \u2014 degraded performance, partial access issues\n\n#### **Summary**\n\nOn July 28, 2026, at 12:47 PM CEST, users began experiencing degraded performance and intermittent difficulty accessing Advanced Insight & Miuros. Our team identified the issue and began investigating shortly after it was detected. A fix was implemented, and full service was confirmed restored by 1:15 PM CEST.\n\n#### **Impact**\n\nBetween 12:47 PM and 1:15 PM CEST, some users experienced slow load times or were unable to reliably access Advanced Insight & Miuros. No data loss occurred, and access was fully restored once the fix was deployed.\n\n#### **Root Cause**\n\nThe degraded performance was traced to a change introduced during a recent deployment to the Advanced Insight & Miuros environment, which had an unintended impact on service performance. Once flagged, our engineering team investigated the affected components and implemented a corrective fix to restore normal performance.\n\n#### **Resolution**\n\nUpon detecting the issue at 12:47 PM CEST, our team began investigating the affected services. The root cause was traced to the recent deployment, and a fix was rolled out, resolving the degraded performance. Full functionality was confirmed at 1:15 PM CEST.\n\n#### **Preventive Measures**\n\nTo help prevent similar incidents going forward, we are:\n\n* Strengthening pre-deployment validation and performance testing for Advanced Insight & Miuros\n* Improving monitoring and alerting thresholds to detect performance degradation earlier\n* Refining our rollback process to reduce time-to-resolution for similar issues\n\n#### **Closing Note**\n\nWe understand how important reliable access to Advanced Insight & Miuros is to your workflows. We apologize for any inconvenience caused during this period and are committed to strengthening the resilience of this service.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-28T13:15:37.739+02:00",
"resolved_inferred": false,
"started_at": "2026-07-28T12:47:20.793+02:00",
"state": "postmortem",
"title": "Advanced Insight & Miuros: Degraded performance",
"updated_at": "2026-08-05T08:45:28.972+02:00",
"url": "https://stspg.io/t9zx7mjmj9fv"
},
{
"body": "### **Post-Mortem: Conversation Handling & Login Service Disruption**\n\n**Date of Incident:** July 28, 2026\n\n**Duration:** 8:53 AM \u2013 9:03 AM CEST \\(10 minutes\\)\n\n**Severity:** High \u2014 partial service disruption\n\n#### **Summary**\n\nOn July 28, 2026, at 8:53 AM CEST, Dixa experienced a service disruption affecting conversation handling and user login functionality. The issue was identified quickly by our engineering team, a fix was deployed at 9:03 AM CEST, and full functionality was restored within 10 minutes of the initial impact.\n\n#### **Impact**\n\nUsers may have experienced difficulty logging into Dixa or handling conversations during the affected window. No data loss occurred, and all systems returned to normal operation by 9:03 AM CEST.\n\n#### **Root Cause**\n\nThe disruption was traced to a change introduced during a routine deployment to our production environment, which had an unintended effect on the services responsible for conversation handling and authentication. Our monitoring systems flagged the anomaly shortly after it occurred, allowing our engineering team to identify and roll back the change quickly.\n\n#### **Resolution**\n\nEngineers on-call were alerted immediately once the anomaly was detected. The team traced the issue to the recent deployment and deployed a corrective fix at 9:03 AM CEST, which resolved the disruption.\n\n#### **Preventive Measures**\n\nTo reduce the likelihood of similar incidents going forward, we are:\n\n* Strengthening our pre-deployment validation checks to catch this class of issue before it reaches production\n* Improving automated rollback triggers so similar anomalies are reverted even faster\n* Enhancing monitoring coverage on conversation handling and authentication services specifically\n\n#### **Closing Note**\n\nWe know reliability is critical to how you run your business, and we take incidents like this seriously. We apologize for any inconvenience this may have caused and remain committed to continuously improving the resilience of our platform.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-28T09:00:00.000+02:00",
"resolved_inferred": false,
"started_at": "2026-07-28T09:00:00.000+02:00",
"state": "postmortem",
"title": "Services not fully functioning",
"updated_at": "2026-08-05T08:44:39.771+02:00",
"url": "https://stspg.io/gr8y278qkj5b"
},
{
"body": "All data is now re-deployed and we're completely up to date for everyone.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "critical",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-06-29T17:11:23.915+02:00",
"resolved_inferred": false,
"started_at": "2026-06-29T11:04:35.000+02:00",
"state": "resolved",
"title": "Advanced Insights unavailable",
"updated_at": "2026-06-29T17:11:23.931+02:00",
"url": "https://stspg.io/rr5lb5z1zhsg"
},
{
"body": "Data ingestion has fully caught up and the Real-Time Dashboard, Search and Conversations pages are showing live data again.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-04-23T13:51:38.603+02:00",
"resolved_inferred": false,
"started_at": "2026-04-23T13:42:51.123+02:00",
"state": "resolved",
"title": "Update delays on Conversation view and Real-Time Dashboard",
"updated_at": "2026-04-23T13:51:38.619+02:00",
"url": "https://stspg.io/bqbbk474qlrk"
},
{
"body": "## **Summary**\n\nOn April 8, 2026, Dixa experienced a partial outage lasting approximately 58 minutes. New WebSocket connections were unable to be established, which meant that new logins failed, and any agents who refreshed their browser or lost their connection could not reconnect. Agents who remained on an existing session were unaffected during the incident.\n\nThe root cause was a TLS certificate misconfiguration introduced during a planned migration of our ingress controller infrastructure. The issue was identified, fixed, and fully resolved within the hour.\n\n## **Impact**\n\nBetween 18:26 and 19:24 CEST, customers attempting to log in to Dixa or re-establish a WebSocket connection \\(e.g., after a page reload\\) were unable to do so. Browsers rejected the connection due to an invalid TLS certificate being served.\n\nAgents who were already logged in with an active WebSocket session continued to operate normally throughout the incident. The impact was limited to new or reconnecting sessions.\n\nA small number of customers were affected and reported the issue to our support team.\n\nNo conversations or data have been lost during the incident.\n\n## **Timeline \\(CEST\\)**\n\n* 18:26 - Internal reports that Dixa is not loading for some users\n* 18:30 - Issue escalated to engineering via our critical support channel\n* 18:32 - Status page updated to **Investigating**\n* 18:40 - Engineering identifies a TLS certificate error \\(`ERR_CERT_AUTHORITY_INVALID`\\)\n* 18:52 - Root cause identified, a self-signed default certificate was being served instead of the correct one\n* 19:00 - Status page updated to **Identified**\n* 19:11 - Fix deployed; WebSocket connections begin recovering\n* 19:12 - Status page updated to **Monitoring**\n* 19:24 - Full recovery confirmed; status page updated to **Resolved**\n\n## **Root Cause**\n\nAs part of our ongoing WebSocket resilience work to ensure a more stable platform, we migrated our ingress controller \\(the component that routes incoming traffic to internal services\\) from an end-of-support solution to a new one. This migration was tested in our staging environment before being applied to production.\n\nHowever, there was a configuration discrepancy between staging and production for the ingress class that handles WebSocket traffic. When DNS was switched to the new ingress controller on the morning of April 8, existing connections continued to work through cached DNS entries still pointing to the old controller. Hours later, as DNS caches expired across the internet, clients began resolving to the new controller, which, due to the misconfiguration, did not recognize the WebSocket routes. This caused it to serve a default self-signed TLS certificate instead of the valid one, leading browsers to reject the connection.\n\nThe length of the partial outage per customer varies depending on when the DNS cache expired, and if WebSocket connections started to connect to the new load balancer.\n\n## **Resolution**\n\nOnce the root cause was identified, a configuration update was deployed to the new ingress controller so that it could correctly handle WebSocket traffic. Connections began recovering immediately after the fix was applied.\n\n## **Preventive Measures**\n\nWe have taken the following steps to reduce the likelihood and impact of similar issues in the future:\n\n* **Environment parity:** All non-production environments have been aligned with production configuration conventions, eliminating the discrepancy that caused this incident.\n* **Endpoint monitoring:** We have added external monitoring checks that validate both the availability and TLS certificate validity of our WebSocket endpoints. This will enable faster detection if a similar issue occurs.\n* **WebSocket isolation**: We will isolate the platform's WebSocket requirement to be optional, so if a similar issue should happen in the future, the disruption will be less intrusive for users.\n\n\u200c\n\n\u200c\n\nWe sincerely apologize for the disruption this caused. Reliability is a top priority for us, and we are committed to learning from every incident to make Dixa more resilient. If you have any questions, please don't hesitate to reach out to your account team or our support at [friends@dixa.com](mailto:friends@dixa.com).",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-04-08T19:24:01.536+02:00",
"resolved_inferred": false,
"started_at": "2026-04-08T18:32:27.000+02:00",
"state": "postmortem",
"title": "Partial outage of Agent Interface connections",
"updated_at": "2026-04-10T14:27:50.210+02:00",
"url": "https://stspg.io/8pwz7dpdm6y8"
},
{
"body": "_Dixa experienced issues with the Smart Reply feature in AI CoPilot on 18 March 2026._  \n  \n_**Summary**_  \n_18 March 2026 at 14:37 CET - 20:20 CET, customers experienced issues with the Smart Reply feature in AI CoPilot. Agents were unable to use Smart Reply._\n\n  \n_**Root cause**_  \n_A code change introduced a regression that caused agents' browsers to block completing the request to load Smart Reply suggestions._  \n  \n_**Timeline**_  \n_At 14:37 CET on 18 March 2026: Issue flagged internally after several customers reported the issue_  \n_At 15:33 CET: Engineering investigates and attempts an initial rollback, which does not resolve the issue._  \n_At 16:02 CET: Status page updated to \"Investigating\" \u2014 customers notified of difficulties with AiCoPilot Smart Replies._  \n_At 16:48 CET: Status page updated \u2014 engineering continues to investigate._  \n_At 18:17 CET: Status page updated \u2014 root cause investigation ongoing_  \n_At 20:20 CET: The knowledge-backend service is successfully rolled back to a stable version, restoring Smart Reply._  \n_At 20:52 CET: Status page updated to \"Resolved\" \u2014 all known issues confirmed resolved._  \n  \n_**Solution**_  \n_The immediate solution was to roll back to a stable version from prior to the breaking change, restoring Smart Reply for all affected customers._  \n_Long term, the team will review the offending commit before re-deploying to prevent similar regressions from reaching production undetected._  \n_We sincerely apologise for the inconvenience caused by this issue._",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-03-18T20:52:32.662+01:00",
"resolved_inferred": false,
"started_at": "2026-03-18T16:02:38.364+01:00",
"state": "postmortem",
"title": "Degraded performance: AiCoPilot: Smart replies",
"updated_at": "2026-03-26T09:57:21.952+01:00",
"url": "https://stspg.io/p4cdnnshd0jw"
},
{
"body": "**Summary:** On March 11, 2026, Dixa experienced platform-wide degraded performance lasting approximately 3 hours \\(08:30 - 12:35 CET\\). Customers experienced slow or failed conversation loading, timeouts on email sending, conversation transfers, assignments, and flow processing. No data was lost, and there were no security issues at any point.\n\n**Impact:**\n\n* _Availability_: Platform-wide slowness and partial inaccessibility for ~3 hours.\n* _Affected functionality_: Conversation loading, email sending, conversation transfers, conversation assignments, and flow processing - all experienced significant slowness and intermittent failures.\n* _Data integrity:_ All emails were fully processed after the fix. No data was lost, and no security issues occurred at any point.\n\n**Root Cause:** The incident was caused by an atypical traffic pattern in inbound email processing that resulted in repeated internal retries - retries are a normal part of email distribution, accounting for factors such as sending delays and server availability, but this expanded exponentially. The sustained retry volume placed excessive load on a central platform component, causing cascading timeouts across dependent services and resulting in platform-wide degradation.\n\n**Timeline \\(CET\\):**\n\n* Mar 11, 06:00 - First signs of email processing errors detected\n* Mar 11, 08:30 - Platform degradation begins; customer impact starts\n* Mar 11, 10:50 - First mitigation deployed; partial improvement\n* Mar 11, 12:30 - Root cause fully identified; final fix applied\n* Mar 11, 12:35 - Platform stability confirmed\n\n**Resolution:** We identified and addressed the source of the abnormal email volume, which resulted in an immediate reduction in error rates, and the platform to recover.\n\n**What We Have Done Since This Incident:** Following this incident, we have already implemented the following improvements:\n\n1. Added validation to reject invalid email addresses early in the pipeline, preventing them from entering retry loops.\n2. Optimised internal lookups to fetch only necessary data instead of the full conversation history, significantly reducing load during email processing.\n3. Added deduplication logic to prevent redundant data fetches during email processing.\n4. Enforced concurrency limits: platform components now shed excess traffic when saturated, allowing requests to be redistributed rather than queued indefinitely.\n5. Added deadline checking: expired requests are now discarded immediately instead of consuming resources on work that is no longer needed.\n6. Reduced internal timeout thresholds to fail fast under contention rather than blocking for extended periods.\n\n**What We're Continuing to Work On:**\n\n1. Loop detection and interruption - Introduce mechanisms to detect and automatically halt email processing anomalies before they can accumulate significant load.\n2. Improved alerting and escalation - Ensure processing anomalies are detected and escalated with appropriate urgency.\n\n**Closing Note:** We sincerely apologize for the disruption this caused. These improvements are our highest priority. If you have any questions, please reach out to [friends@dixa.com](mailto:friends@dixa.com).",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-03-11T12:31:08.544+01:00",
"resolved_inferred": false,
"started_at": "2026-03-11T09:18:20.141+01:00",
"state": "postmortem",
"title": "Degraded performance",
"updated_at": "2026-03-13T13:27:43.867+01:00",
"url": "https://stspg.io/9zb167xystfd"
},
{
"body": "**Summary**: On March 10, 2026, all public Dixa Knowledge Help Centers were unavailable following a production deployment performed during scheduled maintenance. \n\n**Impact**\n\n* Dixa Knowledge Public Help Centers remained inaccessible for 6 hours \\(approximately\\) \n* Customers with advanced custom CSS styling were contacted directly with instructions on how to update their selectors to prevent recurrence.\n\n**Root Cause:** After the production deployment performed during a scheduled maintenance, all public Help Centers became unavailable. Visitors saw an error page \\(500\\) instead of help content.\n\n**Timeline \\(CET\\)**\n\nMar 10, 08:00 - Maintenance completed. No anomalies detected.\n\nMar 10, 11:52 - Initial customer reports received. Considered isolated cases at the time.\n\nMar 10, 13:38 - Incident process triggered. Reported on Status page.\n\nMar 10, 13:57 - Root cause fully identified.\n\nMar 10, 14:13 - Fix deployed. Public Help Centers restored; some styling issues persisted.\n\nMar 10, 14:48 - Root cause for styling anomalies identified as linked to build-time generated class names. Affected customers notified with instructions.\n\nMar 10, 18:03  - Incident resolved. \n\n**Resolution**: We identified a misconfiguration in the deployment that caused Help Centers to attempt to load from a test environment rather than production. A fix was deployed, and an immediate reduction in error rates confirmed it was effective.\n\n**What We Have Done Since This Incident:** Following this incident, we have already implemented the following improvements:\n\n1. Proactively contacted all affected customers with clear instructions on how to update their custom CSS styling.\n2. Documented additional validations for future migrations and action plans to address similar errors. \n3. External documentation updated to offer alternatives to the use of auto-generated classes for styling Help Centers.\n\n**What We're Continuing to Work On**\n\n* Improve monitoring and alerting, ensure availability issues are detected and escalated promptly, independent of customer reports.\n* Investigating a solution to expose stable, named elements for Help Center customization, reducing dependency on build-generated class names.\n\n**Closing Note:** We sincerely apologize for the disruption this caused. We appreciate your patience throughout this disruption. If you have any questions, please reach out to friends@dixa.com.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-03-10T18:03:48.756+01:00",
"resolved_inferred": false,
"started_at": "2026-03-10T13:38:16.571+01:00",
"state": "postmortem",
"title": "Degraded performance",
"updated_at": "2026-03-16T18:04:21.262+01:00",
"url": "https://stspg.io/pfn5sz4fd6r3"
},
{
"body": "**Summary**  \nOn March 2, 2026, Dixa experienced platform-wide degraded performance lasting approximately 3 hours and 45 minutes. Customers experienced slow or failed conversation loading, timeouts on email sending, conversation transfers, assignments, and flow processing. No data was lost, and there were no security issues at any point.\n\n**Impact**\n\n* **Availability:** Platform-wide slowness and partial inaccessibility for ~3 hours 45 minutes.\n* **Affected functionality:** Inbound email processing, conversation loading, conversation transfers, conversation assignments, and flow processing \u2014 all experienced significant slowness and intermittent failures.\n* **Side effect:** Some inbound emails resulted in empty, queue-less conversations being created. These are safe to close or merge with the correctly processed follow-up conversation.\n* **Data integrity:** All emails were fully processed after the fix. No data was lost and no security issues occurred at any point.\n\n**Root Cause**  \nThe incident was caused by a significant and atypical spike in inbound conversations that fell well outside expected operational parameters. This unexpected volume put pressure on a central database component, which began throttling requests. The throttling then cascaded to other parts of the system, resulting in platform-wide slowness and temporary inaccessibility.\n\n**Timeline \\(CET\\)**  \nMar 1, 02:14 Significant and atypical spike in inbound conversations begins  \nMar 2, ~17:45 Database throttling begins; connection pools saturate  \nMar 2, 19:08 Incident declared; investigation begins   \nMar 2, 20:06Mitigation measures applied; investigation ongoing  \nMar 2, 20:18 Root cause identified  \nMar 2, 21:09 Fix deployed; error rates drop  \nMar 2, 21:17 Resolution confirmed  \nMar 2, 21:21Incident resolved\n\n**Resolution**  \nWe identified and addressed the source of the abnormal conversation volume, which immediately caused error rates to drop and the platform to recover.\n\n**What We're Doing to Prevent Recurrence**  \nWe have identified several systemic improvements and are actively working on them:\n\n1. **Detect and suppress atypical inbound volume patterns** \u2014 Introduce early filtering to prevent abnormal spikes from reaching core platform components.\n2. **Improve retry and backoff behaviour** \u2014 Reduce the risk of compounding load during high-traffic failure scenarios.\n3. **Reduce inter-service dependencies** \u2014 Limit the blast radius of a single overloaded component affecting other parts of the platform.\n4. **Improve database resilience** \u2014 Add circuit breakers and timeouts to prevent database pressure from cascading across services.\n5. **Atomic conversation creation** \u2014 Ensure conversations are never persisted without their initial message, eliminating orphaned empty conversations as a failure side effect.\n\n**Closing Note**  \nWe sincerely apologize for the disruption this caused. We take platform reliability seriously and are committed to the systemic improvements outlined above. If you have any questions, please reach out to [friends@dixa.com](mailto:friends@dixa.com).",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-03-02T21:21:46.989+01:00",
"resolved_inferred": false,
"started_at": "2026-03-02T19:08:07.114+01:00",
"state": "postmortem",
"title": "Degraded performance",
"updated_at": "2026-03-04T16:26:50.335+01:00",
"url": "https://stspg.io/4655y75j34qq"
},
{
"body": "**Incident Summary**  \nOn Feb 20, 2026 - 15:24 CET, we experienced a brief period of degraded performance across the platform. Some customers may have noticed slower response times and intermittent timeouts while using the interface.\n\n**Impact**  \nDuring the incident window, a subset of requests were delayed or failed, which may have affected normal usage for some customers.\n\n**Root Cause**  \nThe issue was caused by an internal service experiencing elevated load, which temporarily impacted request processing across the platform.\n\n**Resolution**  \nOur engineering team identified the issue quickly and restored normal performance. The platform has been operating normally since the fix was applied.\n\n**Prevention**  \nWe are reviewing monitoring and scaling controls for the affected service to reduce the likelihood of similar incidents in the future.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-02-20T15:35:23.000+01:00",
"resolved_inferred": false,
"started_at": "2026-02-20T15:24:47.000+01:00",
"state": "postmortem",
"title": "Degraded performance",
"updated_at": "2026-02-25T08:06:55.773+01:00",
"url": "https://stspg.io/qv0n5f86x4sc"
},
{
"body": "Dixa experienced issues with transcription on Tuesday 27th January 2026 between 11:27 and 13:15 CET\n\n  \n**Timeline**\n\n* At 10:22 CET / 4:22 EST Azure reports issues on the Azure OpenAI Service \n* At 11:27 CET / 5:27 EST the first issues with transcriptions start to happen at Dixa. At 11:46 CET / 5:46 EST Dixa investigates unrelated issues with live chats and calls. \n* At 12:16 CET / 6:16 EST While investigating the unrelated issues, Dixa noticed the issues with transcriptions. \n* At 13:15 CET / 7:15 EST Transcription issues disappear.\n\n**Impact**  \nTranscriptions were not available or delayed in the specified timeframe.  \n  \n**Root cause**   \nOur provider for AI voice transcriptions, Azure, reported issues. According to their [status](https://azure.status.microsoft/en-us/status/history/):  \n_Between 09:22 UTC and 16:12 UTC on 27 January 2026, and again between 11:14 UTC and 13:35 UTC on 29 January 2026, a platform issue resulted in an impact to the Azure OpenAI Service in the Sweden Central region. Impacted customers may have experienced HTTP 500/503 errors, failed inference requests, and/or issues with model deployment metadata. This issue also affected the Agent Service and other downstream AI Services dependent on Azure OpenAI in this region._  \n  \n**Immediate solution**  \nWe monitored the situation until it was fixed.  \n  \n**Long-term solution \\(where applicable\\)**  \nWe could consider changing region for our service if the issues become more frequent. Otherwise, Azure has been reliable enough to keep the current infrastructure. Furthermore, Azure made improvements on their side to prevent this situation from happening again.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-01-27T13:36:22.295+01:00",
"resolved_inferred": false,
"started_at": "2026-01-27T12:41:10.851+01:00",
"state": "postmortem",
"title": "Degraded performance - AI Voice transcript",
"updated_at": "2026-02-02T13:40:37.351+01:00",
"url": "https://stspg.io/q674hr1z88t0"
},
{
"body": "All known issues to this incident have been resolved. We thank you for your patience and cooperation.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-01-06T16:40:16.746+01:00",
"resolved_inferred": false,
"started_at": "2026-01-06T16:24:58.794+01:00",
"state": "resolved",
"title": "Degraded performance - Contact Search in Outbound Calling",
"updated_at": "2026-01-06T16:40:16.764+01:00",
"url": "https://stspg.io/9qh8r0zw7t0v"
},
{
"body": "# **Degraded performance - AI Voice Transcription and webhooks**\n\n## **Summary**\n\nOn 22 December 2025, Dixa experienced an issue with the AI Voice Transcription feature. Transcriptions were not being generated for phone calls. Once the issue was resolved, a temporary period of degraded performance on Webhooks occurred as the system processed a backlog of pending transcriptions.\n\n## **Root cause**\n\nFollowing a scheduled release, a component responsible for processing voice transcription requests stopped functioning correctly. When the issue was resolved, the system began processing all pending transcription requests. The high volume of these requests temporarily affected Webhook delivery times, causing delays in notifications to third-party integrations.\n\n## **Timeline**\n\n**Monday 22 December at 11:30 CET:** Issue reported and investigation started.\n\n**Monday 22 December at 12:00 CET:** The issue was identified and a fix was deployed. AI Voice Transcription began processing again.\n\n**Monday 22 December at 12:26 CET:** Degraded Webhook performance identified, affecting some third-party integrations, including chatbots.\n\n**Monday 22 December at 12:54 CET:** AI Voice Transcription fully restored.\n\n**Monday 22 December at 14:35 CET:** All services fully recovered. Incident closed.\n\n## **Preventive measures**\n\nDixa will implement the following measures to prevent similar issues in the future:\n\n* Enhanced monitoring for the AI voice transcription feature.\n\n* Improved handling of backlog processing to minimise impact on dependent services.\n\nWe sincerely apologise for the inconvenience this has caused.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-12-22T14:41:38.098+01:00",
"resolved_inferred": false,
"started_at": "2025-12-22T11:30:19.148+01:00",
"state": "postmortem",
"title": "Degraded performance - AI voice transcriptions and delayed webhook events",
"updated_at": "2026-01-13T12:48:12.645+01:00",
"url": "https://stspg.io/b0pr4vyhxvbc"
},
{
"body": "**Summary**\n\nOn the 16th of December 2025 between 8:49 PM CET and 9:34 PM CET we experienced intermittent issues with AI features such as Co-pilot and intent detection automation on some of the Dixa instances.\n\n**Root cause**\n\nThe incident was caused by the increased usage of the above mentioned features which caused hitting the set usage limit. \n\n**Timeline**\n\nAt 8:49 PM CET: We started observing usage spikes which resulted in errors that started occurring around AI related features. \n\nAt 8:55 PM CET: Status page was updated with the areas that were impacted by the issue.\n\nAt 9:03 PM CET: We found the root cause and moved the incident status to Identified.  \n\nAt 9:12 PM CET: The fix was deployed and the incident status changed to Monitoring.\n\nAt 9:34 PM CET: After a period of monitoring we did not observed any more errors and the incident status was moved to Resolved. \n\n**Solution**\n\nThe immediate solution was to increase the usage limits as well as implement the separation of the limits per each AI feature in the backend. \n\nLong term we are planning to implement proactive alerting and monitoring around AI features usage and set limits to prevent the errors that occurred when the usage spiked. \u200c\n\nWe sincerely apologise for the inconvenience caused by the above issue.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-12-16T21:34:11.069+01:00",
"resolved_inferred": false,
"started_at": "2025-12-16T20:49:59.000+01:00",
"state": "postmortem",
"title": "Intermittent issues with AI features",
"updated_at": "2025-12-23T17:51:24.020+01:00",
"url": "https://stspg.io/zmz0d9479dck"
},
{
"body": "Summary\nParts of Dixa\u2019s search functionality malfunctioned during a 30-minute window, following a deployment including some deficient code, affecting, amongst others, the search feature within the Email and Phone composer features, respectively. \n\nRoot Cause\nStarting at 2.28 PM (CET) on 8 Dec, the faulty application started serving requests, some of which it was unable to handle. This was due to an insufficiently tested code path that got invoked under certain feature flag configurations, resulting in the search request not being handled correctly.\n\nAction Items\nOn-call engineers were alerted about elevated error levels within 15 minutes of the first failing requests, and shortly after, were able to identify the impaired deployment, rolling it back to a previous version while investigations continued. At 3.07 PM (CET), error rates related to failing search requests had completely levelled off.\n\nA fix to the root issue has been released since, further improving the overall safety and handling of code dictated by feature flag configurations. This includes more thorough runtime parameter checks as well as adding several regression tests to ensure the issue won\u2019t be reintroduced.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-12-08T14:30:00.000+01:00",
"resolved_inferred": false,
"started_at": "2025-12-08T14:30:00.000+01:00",
"state": "resolved",
"title": "Search - Partial Outage",
"updated_at": "2025-12-11T07:52:11.304+01:00",
"url": "https://stspg.io/69tg8fwlp36t"
},
{
"body": "**Summary** \n\nOn the 18th of November 2025 at 12:55 PM CET we received reports of Dixa Knowledge help centers not being available due to a Cloudflare error. At roughly the same time, we were made aware internally that we were receiving a lot of errors on AI-related functionality from one of our upstream sub-processors handling AI requests.\u200c\n\n**Root cause** \n\nThe incident originated from a major Cloudflare service disruption \\([https://www.cloudflarestatus.com/incidents/8gmgl950y3h7\\)](https://www.cloudflarestatus.com/incidents/8gmgl950y3h7)) caused by a configuration file used for Cloudflare\u2019s Bot Management system. The file was auto-generated and due to a bug became too large, subsequently crashing Cloudflare services and causing major disruption across the globe. \n\nDixa is using Cloudflare solely for the purpose of serving Dixa Knowledge's Knowledge Bases. Cloudflare helps us handle verification of custom host names and generating certificates for these custom host names, as well as protect the knowledge bases from malicious third parties.\n\nWe quickly learned that one of our upstream sub-processors for AI-related features also uses Cloudflare for similar reasons, therefore also breaking AI Co-Pilot, Mim, Intent Detection and other AI-powered features inside Dixa.\n\n**Timeline**\n\nAt 12:55 pm CET: We started observing increased errors on services handling AI-related features inside Dixa and received the first report about Dixa Knowledge KBs being inaccessible. We triggered our internal incident process and communicated here at 12:59 PM CET.\n\nAt 1:06 pm CET: The full impact on Dixa services started to become clear, and we were confident only AI-related features and Knowledge Bases were affected. \n\nFor the next hour and a half, we monitored the situation closely and saw alternating service recovery and disruption on the before-mentioned services while Cloudflare was recovering.\n\nAt 2:49 pm CET: We were made aware by our AI sub-processor that they saw slight improvements in connectivity.\n\nAt 3:46 pm CET: The impacted systems have fully recovered. The incident status was moved to Monitoring. \n\nAt 5.43 pm CET: We changed the incident on our status page from 'Monitoring' to 'Resolved' as no further problems have been reported nor detected.\n\nAt 6:44 pm CET: Cloudflare confirmed that the incident is fully resolved on their end.\n\n\u200c\n\nWe sincerely apologize for the inconvenience this has caused.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-11-18T16:43:10.576+01:00",
"resolved_inferred": false,
"started_at": "2025-11-18T12:59:22.233+01:00",
"state": "postmortem",
"title": "Degraded performance",
"updated_at": "2025-11-20T09:23:32.483+01:00",
"url": "https://stspg.io/p9fs6swndf5k"
},
{
"body": "**Summary**\u00a0\n\nOn the 20th of October 2025 at 6:55 PM CEST due to the ongoing AWS incident Dixa experienced intermittent issues with Telephony Inbound and Outbound connectivity. On some Dixa instances the calls were not being placed or the connection between agents and users was not established.\u00a0\n\n\u200c\n\n**Root cause**\u00a0\n\nThe incident originated from a **major AWS service disruption in the us-east-1 region** caused by a **DNS race condition within Amazon DynamoDB\u2019s endpoint management system**. The fault temporarily removed valid DNS records for several AWS services, resulting in widespread API resolution failures across multiple dependent systems.\n\nDixa is using Twilio\u2019s APIs for **telephony services.** When DynamoDB\u2019s DNS records were invalidated, Twilio\u2019s **regional load balancers lost access to key internal routing data**, which caused inbound and outbound voice requests to fail.\u00a0\n\nBecause Twilio\u2019s APIs are globally routed through Twilio\u2019s us-east-1 infrastructure, Dixa\u2019s platform became unable to establish or maintain voice sessions.\u00a0\n\n\u200c\n\n**Timeline**\n\nAt 06:55 pm CEST: We started observing intermittent interruptions with Telephony Inbound and Outbound service. Only some Dixa instances and phone numbers were affected. The incident was reported with the Identified status.\u00a0\n\nAt 07:52 pm CEST: The update is posted stating that some of the services are still impacted.\u00a0\n\nAt 08:51 pm CEST: There are no signs of issues on Dixa services anymore, however the third party services are still recovering and minimal interruptions could be expected.\u00a0\n\nAt 09:46 pm CEST: The impacted systems have fully recovered. The incident status was moved to Monitoring.\u00a0\n\nAt 00:56 am CEST on the 21st of October: AWS confirmed that the incident is fully resolved on their end. Dixa\u2019s incident status was changed to Resolved.\u00a0\n\n\u200c\n\n**Preventive measures**\u00a0\n\nDixa will evaluate implementing the following measures to mitigate the risk of similar issues in the future:\u00a0\n\n* evaluate the feasibility of enabling outbound and inbound routing via multi-regional providers in case the us-east-1 Twilio region fails.\n* Independent external health check monitoring outside of AWS and Twilio.\n* Service dependency mapping and redundancy revisit.\n* Incident communication redundancy\u00a0\n\n\u200c\n\nWe sincerely apologise for the inconvenience this has caused.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-10-21T00:56:58.770+02:00",
"resolved_inferred": false,
"started_at": "2025-10-20T18:55:44.806+02:00",
"state": "postmortem",
"title": "Degraded performance - telephony services",
"updated_at": "2025-10-28T16:46:39.679+01:00",
"url": "https://stspg.io/6gyk0t4b8hz3"
},
{
"body": "**Summary**\u00a0\n\nOn the 20th of October 2025, starting from 08:58 am CEST, Dixa experienced widespread disruptions caused by the issues coming from third-party services.\u00a0\n\n**Impacted services**\n\n* Telephony Inbound and Outbound traffic - Calls were not being placed for the most part of the incident, and for a period of time, calls were placed, but the connection between agents and end users was not established.\u00a0\n* WhatsApp channel - conversations were queued but not sent during the incident. All queued messages were sent as soon as the issue was resolved.\n* SMS channel - conversations were queued but not sent during the incident. All queued messages were sent as soon as the issue was resolved.\n* Login to Elevio - login to Elevio was impacted. Help Centers were fully operational.\u00a0\n* Login to the Status page was impacted, which caused challenges with communicating incident updates once the incident was raised.\u00a0\n* Access to Dixa\u2019s Public API documentation[ docs.dixa.io](http://docs.dixa.io) was also affected by the incident; Dixa\u2019s API was fully operational.\n\n**Root cause**\u00a0\n\nThe incident originated from a **major AWS service disruption in the us-east-1 region** caused by a **DNS race condition within Amazon DynamoDB\u2019s endpoint management system**. The fault temporarily removed valid DNS records for several AWS services, resulting in widespread API resolution failures across multiple dependent systems.\n\nDixa is using Twilio\u2019s APIs for **telephony, WhatsApp, and SMS channels**. When DynamoDB\u2019s DNS records were invalidated, Twilio\u2019s **regional load balancers lost access to key internal routing data**, which caused inbound and outbound voice requests to fail. This also impacted dependent APIs for SMS and WhatsApp delivery.\n\nBecause Twilio\u2019s APIs are globally routed through Twilio\u2019s us-east-1 infrastructure, Dixa\u2019s platform became unable to establish or maintain voice sessions, send or receive SMS messages, or process WhatsApp traffic. Other parts of the Dixa platform \\(login, chat, email\\) remained operational, but all Twilio-dependent communication channels were impacted for the most part of the duration of the AWS event.\u00a0\n\n### **Timeline**\n\nAt 08:58 am CEST: We started observing errors in the logs coming from 3rd party systems, and shortly after, we started getting reports from customers about issues.\n\nAt 09:16 am CEST: The Incident was reported on the status page with status: Investigating. The system we use for the status page was also impacted due to the same root cause at AWS, and we were unable to add further details regarding the incident.\n\nBetween 9:00 and 10:00 am CEST: We kept updating the users who reached out to our support team about the status of the issue, while adding more details to the status page was not possible.\n\nAt 10:00 am CEST: An updated incident message was posted in the Agent interface to notify users about the status of the incident\u00a0\n\nAt 10:07 am CEST: We regained access to the status page and updated the status of the incident to: Identified and added the information about impacted services.\u00a0\n\nAt 10:20 am CEST: The incident status update is posted as well, with the notification being sent to all status page subscribers.\u00a0\n\nAt 10:55 am CEST: The next status page update is posted with reference to the incidents at AWS and Twilio, which were the root cause of the issues with Dixa services. We also reported issues with logging in to Elevio caused by the same incident.\u00a0\n\nAt 11:52 am CEST: Update is posted. The Elevio login issue is resolved. The issue with telephony inbound and outbound, as well as SMS and WhatsApp service, continues.\u00a0\n\nAt 12:58 pm CEST: The first calls were starting to get to Dixa infrastructure again, but the voice in the calls was still not operational.\u00a0\n\nAt 13:05 pm CEST: The connection issue with telephony continues. SMS is confirmed to be fully operational again.\u00a0\n\nAt 13:26 pm CEST: The inbound and outbound telephony is confirmed to be operational again.\n\nAt 14:21 pm CEST: WhatsApp is confirmed to be operational as well. All messages that were queued during the incident got successfully delivered. The incident status is changed to: Monitoring\n\nAt 15:19 pm CEST: All systems have been operational since the incident was moved to Monitoring, and the incident status was changed to: Resolved.\u00a0\n\n\u200c\n\n**Preventive measures**\u00a0\n\nDixa will evaluate implementing the following measures to mitigate the risk of similar issues in the future:\u00a0\n\n* evaluate the feasibility of enabling outbound and inbound routing via multi-regional providers in case the us-east-1 Twilio region fails.\n* Independent external health check monitoring outside of AWS and Twilio.\n* Service dependency mapping and redundancy revisit.\n* Incident communication redundancy\u00a0\n\n\u200c\n\nWe sincerely apologise for the inconvenience this has caused.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-10-20T15:19:26.640+02:00",
"resolved_inferred": false,
"started_at": "2025-10-20T09:16:46.353+02:00",
"state": "postmortem",
"title": "Multiple services disruptions",
"updated_at": "2025-10-28T15:05:02.760+01:00",
"url": "https://stspg.io/m8cjfld1zvj6"
},
{
"body": "### Summary\n\nDixa experienced partial degraded performance of AI-Co pilot features on October 14 2025 between 08:13-13:13 CEST. Users who never changed their translation settings were affected, experiencing unwanted auto-translations and non working translations. The issues were cause by an increased number of translation requests hitting our api, leading to overload. We eventually solved it by increasing the capacity for handing requests and by reverting the changes  \n  \n**Timeline**\n\nAt 08:13 am CEST A change was deployed for AI copilot to default the Auto summary and translations  to \u201cAutomatic\u201d for users who never changed their preferences\n\nAt 09:48 am CEST First reports of difficulties with translations for AI Co-pilot that didn\u2019t work at all.\n\nAt 10:23 am CEST A fix is deployed to support the increased amount of requests incoming.\n\nAt 10:33 am CEST Fix is evaluated to have mitigated the issues.\n\nAt 11:33 CEST The requests to our API increase and issues come back, more reports are coming in from customers about conversations being automatically translated or translations not working at all.\n\nAt ~12:45 am CEST Fix is applied to support even more increased amount of requests to prevent translation failures and the quota on the model used by translation is increased to support the increased amount of requests.\n\nAt 13:13 am CEST A final fix is deployed - changes to AI copilot settings defaults are reverted and requires agents to reload to take effect. This decreased the system overload.\n\n### **Impact**\n\nThe incident had a major impact on the users of the AI Co-pilot feature that never changed their settings who experienced degraded performance with conversations not translating or unwanted increased amount of translations performed.\n\n## **Root cause**\n\nThe frontend changes to settings defaults for AI auto summary and translate flipped the settings from on-demand to automatic. It was applied to existing customers that never changed their settings. It appeared to be a significant amount of customers who had not done any changes to their settings. The change took effect over the course of multiple hours, because it required agents to reload the UI for the settings to take effect. This caused many more translations being performed where customers did not expect. Additionally, this caused increased amount of translation requests via rest-api, which made rest-api incapable of performing all requests due to it having a capped number on max open requests\n\n### **Resolution**\n\n* Support for and quota of translation requests were increased\n* The frontend change to settings defaults for AI Co-pilot were reverted\n\nWe have taken steps to improve our processes for control and increased the capacity limit so that this won\u2019t happen again.  \n  \nWe apologize for the inconvenience this caused and thank you for you cooperation",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-10-14T14:41:17.239+02:00",
"resolved_inferred": false,
"started_at": "2025-10-14T10:24:45.881+02:00",
"state": "postmortem",
"title": "Degraded performance - Translations not working",
"updated_at": "2025-10-17T16:18:03.166+02:00",
"url": "https://stspg.io/gr1d3qdbx04f"
},
{
"body": "## Summary\n\nOn the 9th of September, 2025, a deployment at 2:45 PM CEST caused an ingestion delay affecting the Intelligence and the Realtime Dashboards. The incident was detected at 2:47 PM CEST by internal alerts as well as reports coming in to our Customer Support. A fix was released at 3:25 PM CEST, and all data was caught up at 3:51 PM CEST. The total incident was 66 minutes, with an impact on customers relying on our real-time dashboard and Intelligence module.\n\n## Timeline\n\n* **2:45 PM CEST** - Deployment released that caused the incident where ingestion of data to Intelligence- and the Realtime Dashboards was delayed\n* **2:47 PM CEST** - Internal alerts notified relevant stakeholders at Dixa around the incident, and a fix was worked on\n* **3:25 PM CEST** - A fix was released, and delayed data started catching up\n* **3:51 PM CEST** - All delayed data had caught up  \n  **Total Duration:** 66 minutes\n\n## Root Cause\n\nThe incident was caused by a deployment that resulted in a delay in the ingestion of data.\n\n## Impact\n\n* Realtime Dashboards\n* Intelligence/Analytics\n\n## Resolution\n\n**Immediate Resolution:**\n\n* The engineering team started working on a fix as soon as internal alarms notified relevant stakeholders\n* Service functionality was restored following the fix, and the delayed data started catching up\n\n**Prevention Measures:**\n\n* Expand test coverage",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-09-09T17:49:12.832+02:00",
"resolved_inferred": false,
"started_at": "2025-09-09T15:45:32.841+02:00",
"state": "postmortem",
"title": "Degraded performance for Analytics and Real-time Dashboard",
"updated_at": "2025-09-15T08:52:42.396+02:00",
"url": "https://stspg.io/t5wcxd74xy7j"
},
{
"body": "**Date**: August 22, 2025    \n**Duration**: 5 days \\(August 22-27, 2025\\)    \n**Severity**: Degraded\n\n## Summary\n\nBetween August 22-27, 2025, our AI assistant \\(Mim\\) experienced degraded response quality due to inconsistent knowledge base indexing failures. The incident was caused by our OpenSearch cluster rejecting document index requests due to memory constraints from oversized embeddings. While Mim continued to operate with partial knowledge data, customers experienced reduced response accuracy for 5 days until the issue was resolved through cluster scaling.\n\n## Timeline\n\n* **August 22, 2025:** First indexing failures begin occurring\n* **August 22-25, 2025:** Indexing failures continue to grow in frequency\n* **August 25, 2025:** Issue escalated by CS team\n* **August 27, 2025 - 18:00 CEST:** Indexing failures stop occurring, and new knowledge bases can be indexed successfully\n* **August 28, 2025:** re-indexed existing Dixa Knowledge and elevio data sources\n* **August 28, 2025 - 12:00:** Full service restoration achieved\n\n## Root Cause\n\nThe system was unable to process knowledge base indexing requests due to insufficient computational resources relative to the data volume and complexity. Contributing factors included:\n\n* Inadequate system capacity for the current workload demands\n* Inefficient resource utilization patterns\n* Suboptimal data processing architecture that doesn't scale effectively with growth\n* Large data structures requiring more system resources than available\n\n## Impact\n\n* **Duration**: 5 days \\(August 22-27, 2025\\)\n* **Service Level**: Degraded \\(not complete outage\\)\n* **User Experience**: Mim responses had reduced accuracy and completeness due to operating with partial knowledge data\n\n## Resolution\n\n### Immediate Actions Taken\n\n1. **System Cleanup**: Removed unused data and optimized storage utilization\n2. **Capacity Scaling**: Increased available computational resources to handle workload demands\n3. **Performance Optimization**: Balanced system performance improvements with operational efficiency\n\n### Long-term Improvements Planned\n\n1. **Architecture Enhancement**:\n\n    * Redesign data processing workflows for improved scalability\n    * Implement more efficient system resource management\n    \n2. **Data Optimization**:\n\n    * Evaluate opportunities to reduce data processing overhead\n    * Optimize data formats for better system performance\n    \n3. **Monitoring Enhancement**:\n\n    * Implement proactive alerting for system capacity issues\n    * Add performance monitoring to prevent resource bottlenecks",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-08-28T08:52:13.801+02:00",
"resolved_inferred": false,
"started_at": "2025-08-26T16:30:36.692+02:00",
"state": "postmortem",
"title": "Mim Data Source Sync Failures",
"updated_at": "2025-08-28T15:46:51.766+02:00",
"url": "https://stspg.io/8pb92m5ctnnq"
},
{
"body": "**Date:** August 26, 2025  \n**Duration:** 25 minutes \\(14:20 - 14:45 CEST\\)  \n**Severity:** High - Complete API unavailability\n\n## Summary\n\nA DNS configuration error during staging environment setup caused a complete outage of our public API at [dev.dixa.io](http://dev.dixa.io) for 25 minutes. All external integrations and services dependent on our API were affected.\n\n## Timeline\n\n* **14:20 CEST** - DNS records for [dev.dixa.io](http://dev.dixa.io) inadvertently deleted during staging setup\n* **14:20 CEST** - API connectivity issues began affecting all external services\n* **14:45 CEST** - DNS configuration restored, services resumed normal operation\n\n## Root Cause\n\nDuring the setup of staging DNS entries for [euw1.dev.dixa.io](http://euw1.dev.dixa.io), the existing DNS records for [dev.dixa.io](http://dev.dixa.io) were accidentally removed instead of being preserved alongside the new staging records. This caused all DNS queries for [dev.dixa.io](http://dev.dixa.io) to return NXDOMAIN responses.\n\n## Impact\n\n* **Duration:** 25 minutes of complete API unavailability\n* **Affected Services:** All external integrations including native integrations, custom cards, flow automations, Zapier, chatbots, and third-party API consumers\n* **User Impact:** Complete inability to access API-dependent functionality\n\n## Resolution\n\nThe DNS configuration was corrected by restoring the proper DNS records for [dev.dixa.io](http://dev.dixa.io), immediately resolving the connectivity issues.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "critical",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-08-26T14:20:00.000+02:00",
"resolved_inferred": false,
"started_at": "2025-08-26T14:20:00.000+02:00",
"state": "postmortem",
"title": "DNS Configuration Issue Impacting API Connectivity",
"updated_at": "2025-08-26T15:54:15.805+02:00",
"url": "https://stspg.io/bg68ybr3d23p"
},
{
"body": "## Summary\n\nOn August 26, 2025, a deployment at 10:32 AM CEST caused a complete outage of the TrustPilot queue service, preventing agents from opening TrustPilot conversations in the agent interface. The incident was detected by Customer Support 20 minutes after deployment and resolved through a rollback at 11:30 AM CEST. Total downtime was 58 minutes, with a high impact on TrustPilot customer service operations.\n\n## Timeline\n\n* **10:32 CEST** - Image deployment initiated\n* **10:52 CEST** - Customer Support \\(CS\\) identified the TrustPilot queue service outage and notified the engineering team\n* **11:30 CEST** - Changes rolled back, service restored  \n  **Total Duration:** 58 minutes\n\n## Root Cause\n\nThe incident was caused by a deployment that contained changes incompatible with the TrustPilot queue service functionality.\n\n## Impact\n\n* TrustPilot Conversations could not be opened in the agent interface\n* Complete disruption of TrustPilot queue operations\n* Agent workflow interruption for TrustPilot-related customer inquiries\n\n## Resolution\n\n**Immediate Resolution:**\n\n* The engineering team executed a rollback of the deployed changes at 11:30 AM CEST\n* Service functionality was restored following the rollback\n* TrustPilot queue operations returned to normal\n\n**Prevention Measures:**\n\n* Expand integration test coverage for all third-party services\n* Implement automated health checks and alerting for critical service dependencies",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "critical",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-08-26T11:52:14.701+02:00",
"resolved_inferred": false,
"started_at": "2025-08-26T10:58:35.516+02:00",
"state": "postmortem",
"title": "TrustPilot Queue Service Outage",
"updated_at": "2025-08-27T12:37:38.044+02:00",
"url": "https://stspg.io/dc3fd2f527q8"
}
]
}