{
"vendor": "Pubnub",
"slug": "pubnub",
"platform": "statuspage",
"status_url": "http://status.pubnub.com",
"last_checked": "2026-09-16T12:28:20Z",
"last_state": "ok",
"history_backfilled": true,
"first_watched": "2026-09-04T07:06:16Z",
"incidents": [
{
"body": "### **Problem Description, Impact, and Resolution**\u00a0\n\nOn Tuesday, August 25th, 2026, at 12:10 UTC, our internal monitoring alerted us to an issue where a very small subset of messages within one availability zone of one region \\(our US East point of presence\\) may not have been immediately delivered to subscribers in that same region. All messages were persisted normally within PubNub Persistence service, **thus no message data was lost.** Traffic to or from any other region or AZ was not affected, and no other PubNub services were affected. No customers reported impact.\n\n### **Root Cause**\n\nDuring planned internal load testing, a group of internal routing servers was removed from service. Due to a configuration flaw, their network addresses were not fully retired and remained cached by our publishing layer.\n\nThis presented no problem for the most part. However, one address was subsequently reassigned to an unrelated internal component that responded normally on the same protocol and port. Because that response appeared successful, the publishing layer received no error, and messages sent to that address were routed incorrectly.\n\nA rolling restart of the affected servers cleared the outdated addresses at 13:25 UTC.\n\n### **Mitigation Steps and Recommended Future Preventative Measures**\u00a0\n\n* Once the affected servers were identified, a rolling restart was performed to clear the outdated routing addresses.\n* Following an investigation that confirmed the root cause, a fix ensuring outdated internal routing addresses are correctly retired was deployed to all publishing servers worldwide on August 25th, 2026 at 22:00 UTC.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-08-25T14:51:38.932Z",
"resolved_inferred": false,
"started_at": "2026-08-25T13:22:48.467Z",
"state": "postmortem",
"title": "Potential for some missed messages for subscribers in US East PoP",
"updated_at": "2026-08-28T17:46:59.796Z",
"url": "https://stspg.io/qctzhs1v265m"
},
{
"body": "### **Problem Description, Impact, and Resolution**\u00a0\n\nAt approximately **13:55 UTC on Aug 14, 2026**, we observed elevated publish errors and message replication failures in our publish/subscribe service, which also caused latency in other PubNub services globally. Customers may have experienced increased publish error rates, delayed or missed message delivery, delayed message persistence, and increased latency for Functions and Events & Actions workflows.\n\n\u200cThe root cause of the incident was an unusually large concentration of global publish traffic that was not limited by our throttling layers. That traffic created resource pressure in the publish and replication layers, increased load on storage systems, and caused downstream processing delays in dependent services. We mitigated the issue by adjusting targeted traffic controls, increasing capacity for affected publish and replication components, and isolating the high-volume traffic pattern to reduce broader platform impact. The issue was resolved at approximately **14:15 UTC on Aug 14, 2026**.\n\n\u200cThis issue occurred because of an issue with our automated controls specific to this exceptional traffic pattern. As a result, the increased load affected multiple services before a mitigation could be fully applied.\u00a0\n\n### **Mitigation Steps and Recommended Future Preventative Measures**\n\nTo prevent a similar issue from occurring in the future, we have isolated the identified high-volume traffic pattern onto dedicated infrastructure and also increased baseline capacity for the affected components across our PoPs.\n\nWe are also further strengthening our traffic detection and management processes for this exceptional load pattern so it will be identified and contained earlier without cascading impact across dependent services.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-08-14T14:29:49.948Z",
"resolved_inferred": false,
"started_at": "2026-08-14T14:29:49.892Z",
"state": "postmortem",
"title": "Global Errors and Failures with Publish and Functions, Delays with Events & Actions",
"updated_at": "2026-08-20T16:58:05.421Z",
"url": "https://stspg.io/wd3pg1rlq378"
},
{
"body": "## Problem Description, Impact, and Resolution\u00a0\n\nAt 19:50 UTC on June 10, 2026, we observed a small fraction of publishes originating from US-EAST-1 failing to replicate to subscribers globally. We removed the degraded publisher pod from service and the issue was resolved at 21:21 UTC on June 10, 2026.\u00a0 The root cause of the incident was triggered by a single process that fell into a degraded state where it continued receiving inbound traffic and passing health checks, but traffic sent outbound from the process was failing at an abnormally high rate. Our automated health check/recovery system did not auto-detect and replace the degraded process because its health check API reported itself as healthy.\u00a0\n\n## Mitigation Steps and Recommended Future Preventative Measures\u00a0\n\nTo prevent a similar issue from occurring in the future, we are improving the data within our health check APIs to return more complete performance metrics over a rolling time window. We are enhancing the issue detection logic to detect more patterns that infer process failure, even if the process itself is reporting as healthy.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-06-10T20:45:41.309Z",
"resolved_inferred": false,
"started_at": "2026-06-10T20:14:36.048Z",
"state": "postmortem",
"title": "Replication failures",
"updated_at": "2026-06-16T00:00:33.039Z",
"url": "https://stspg.io/vj0zdnglj3bn"
},
{
"body": "### **Problem Description, Impact, and Resolution**\u00a0\n\nAt approximately **13:20 UTC on June 9, 2026**, we observed elevated publish errors and message replication failures in our publish/subscribe service, which also caused latency in other PubNub services globally. Customers may have experienced increased publish error rates, delayed or missed message delivery, delayed message persistence, and increased latency for Functions and Events & Actions workflows.\n\n\u200c\n\nThe root cause of the incident was an unusually large concentration of global publish traffic that was not limited by our throttling layers. That traffic created resource pressure in the publish and replication layers, increased load on storage systems, and caused downstream processing delays in dependent services. We mitigated the issue by adjusting targeted traffic controls, increasing capacity for affected publish and replication components, and isolating the high-volume traffic pattern to reduce broader platform impact. The issue was resolved at approximately **13:48 UTC on June 9, 2026**.\n\n\u200c\n\nThis issue occurred because we did not have sufficient automated controls and isolation processes in place to protect shared infrastructure from this type of exceptional traffic pattern. As a result, the increased load affected multiple services before mitigation could be fully applied.\u00a0\n\n\u200c\n\n### **Mitigation Steps and Recommended Future Preventative Measures** \n\nTo prevent a similar issue from occurring in the future, we have isolated the identified high-volume traffic pattern onto dedicated infrastructure and also increased baseline capacity for the affected components across our PoPs.\n\n\u200c\n\nWe are also further strengthening our traffic detection and management processes for exceptional load patterns so they can be identified and contained earlier without cascading impact across dependent services.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-06-09T14:00:54.000Z",
"resolved_inferred": false,
"started_at": "2026-06-09T14:00:54.000Z",
"state": "postmortem",
"title": "Global Errors and Failures with Publish and Functions, Delays with Events & Actions",
"updated_at": "2026-06-12T18:33:26.777Z",
"url": "https://stspg.io/pb1d9hy71tpp"
},
{
"body": "### **Problem Description, Impact, and Resolution**\u00a0\n\nOn March 24, 2026, at 19:27 UTC, one network shard experienced intermittent connectivity affecting a subset of customers. The affected users may have experienced elevated latency and temporary error responses related to their subscription requests. The instability was caused by an atypical surge in message volume within a shared processing environment that had improperly configured resource limits. This led to high resource utilization and triggered automated system restarts. PubNub Engineering resolved the issue by implementing the proper limits after expanding infrastructure capacity to accommodate the increased load. Service was fully stabilized once the environment was tuned to the new traffic profile.\n\n### **Mitigation Steps and Recommended Future Preventative Measures**\u00a0\n\n**Infrastructure Tuning:** Adjusted automated scaling parameters to provide greater headroom for rapid traffic fluctuations.\n\n**Enhanced Traffic Management:** Deployed refined monitoring heuristics to better isolate and manage high-volume traffic patterns without impacting shared resources.\n\n**Dynamic Resource Allocation:** Accelerating the rollout of enhanced vertical scaling technology to allow individual processing nodes to adapt more fluidly to demand spikes.\n\n**Operational Coordination:** Strengthening internal protocols for high-capacity events to ensure large-scale traffic shifts are proactively transitioned to dedicated environments.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-03-24T21:12:58.242Z",
"resolved_inferred": false,
"started_at": "2026-03-24T20:21:07.000Z",
"state": "postmortem",
"title": "Connectivity Issues Affecting a Subset of Subscriptions",
"updated_at": "2026-03-25T22:36:41.728Z",
"url": "https://stspg.io/3gr2wjq19vph"
},
{
"body": "### **Problem Description, Impact, and Resolution**\u00a0\n\nOn January 1, 2026 at 00:00 UTC, we observed elevated latency in our History service across multiple regions. Customers may have experienced delays in message persistence and history availability during this period.\n\nThe issue was caused by a mismatch in newly created persistence tables. Specifically, required columns for message metadata were missing from the new tables, resulting in failed write operations and backed-up queues. This created downstream pressure on our storage systems, leading to higher latency in history processing.\n\nWe mitigated the issue by manually applying the correct updates across all affected persistence spaces. After the updates were applied, message processing returned to normal and queue latency cleared.\n\nThis issue occurred because we did not have proper controls in place to ensure schema consistency for newly generated monthly persistence tables.\n\n### **Mitigation Steps and Recommended Future Preventative Measures**\u00a0\n\nTo resolve the issue, we manually applied the required schema updates globally. In the coming days, we will update our change management processes to ensure schema changes are correctly applied to all future monthly tables. We are also auditing our schema tracking and automating validation to prevent inconsistencies across environments.\n\nThese improvements will ensure that future table generation includes all necessary columns and reduce the risk of similar issues impacting History service performance.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-01-01T01:00:33.113Z",
"resolved_inferred": false,
"started_at": "2026-01-01T00:25:52.918Z",
"state": "postmortem",
"title": "Delay in Publishing Messages to Storage Globally",
"updated_at": "2026-01-05T17:03:52.053Z",
"url": "https://stspg.io/x3ssg0swk2d9"
},
{
"body": "**Problem Description, Impact, and Resolution**\n\nStarting at 16:46 UTC on Nov. 20, 2025, we noticed a small number of errors with the publish API in the North American and Asia Pacific regions. The system automatically recovered with all functionality fully restored by 16:50 UTC on Nov. 20, 2025.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-11-20T16:30:00.000Z",
"resolved_inferred": false,
"started_at": "2025-11-20T16:30:00.000Z",
"state": "postmortem",
"title": "Increased errors observed and resolved",
"updated_at": "2025-11-25T01:30:21.411Z",
"url": "https://stspg.io/lg4yq11kq4zn"
},
{
"body": "### **Problem Description, Impact, and Resolution**\u00a0\n\nOn October 20th, 2025 at 07:06 UTC, our monitoring systems alerted us to elevated error levels across multiple PubNub services in the IAD region \\(US-East\\). Some customers may have experienced increased error rates and latency, as well as intermittent issues with Presence service availability across IAD \\(US-East\\), SJC \\(US-West\\), and HND \\(AP-Northeast\\).\n\nWe quickly determined the issue was caused by a broader infrastructure outage affecting our cloud provider \\(AWS\\) in the IAD region. We initiated regional failover procedures and re-routed new connections to alternate regions. However, due to undefined steps in some of our failover processes and delays accessing some tools due to the provider issue, existing connections for some services remained degraded for longer than expected.\n\nTo restore full service, we manually reset established connections, re-routed Presence traffic to Frankfurt \\(EU-Central\\), and brought on additional infrastructure in other regions to absorb traffic. Errors were mitigated by 09:20 UTC. Later in the day, additional regional load in US-West triggered a new wave of service degradation. We responded by isolating the US-East region again and scaling up balancer capacity in US-West. PubNub services were stabilized by 13:20 UTC, and remained in a monitoring state while our infrastructure provider worked to fully resolve the underlying issue.\n\nBy 22:35 UTC, our provider reported full restoration of service. After validating stability in US-East, we completed rebalancing traffic by 23:48 UTC, and declared the incident resolved.\n\n### **Mitigation Steps and Recommended Future Preventative Measures**\u00a0\n\nWhile this incident was caused by an external infrastructure outage, we\u2019ve identified several opportunities to strengthen our internal readiness and response procedures.\n\nWe are consolidating and centralizing our regional failover procedures to ensure they are immediately accessible and complete for all production services. Any gaps in our process documentation for newer services will be addressed to ensure readiness before they are fully adopted into production. Additionally, we are reviewing and resolving issues with internal tooling, including inventory and DNS resolution problems, which made mitigation more difficult during the incident.\n\nThese improvements will ensure faster and more consistent responses to future infrastructure-level disruptions, and reduce potential impact on customer traffic across regions.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-10-20T23:48:53.332Z",
"resolved_inferred": false,
"started_at": "2025-10-20T13:37:33.401Z",
"state": "postmortem",
"title": "Increased latency and errors observed in US-West",
"updated_at": "2025-10-27T18:33:38.262Z",
"url": "https://stspg.io/gy02hdx72zgs"
},
{
"body": "### **Problem Description, Impact, and Resolution**\u00a0\n\nOn October 20th, 2025 at 07:06 UTC, our monitoring systems alerted us to elevated error levels across multiple PubNub services in the IAD region \\(US-East\\). Some customers may have experienced increased error rates and latency, as well as intermittent issues with Presence service availability across IAD \\(US-East\\), SJC \\(US-West\\), and HND \\(AP-Northeast\\).\n\nWe quickly determined the issue was caused by a broader infrastructure outage affecting our cloud provider \\(AWS\\) in the IAD region. We initiated regional failover procedures and re-routed new connections to alternate regions. However, due to undefined steps in some of our failover processes and delays accessing some tools due to the provider issue, existing connections for some services remained degraded for longer than expected.\n\nTo restore full service, we manually reset established connections, re-routed Presence traffic to Frankfurt \\(EU-Central\\), and brought on additional infrastructure in other regions to absorb traffic. Errors were mitigated by 09:20 UTC. Later in the day, additional regional load in US-West triggered a new wave of service degradation. We responded by isolating the US-East region again and scaling up balancer capacity in US-West. PubNub services were stabilized by 13:20 UTC, and remained in a monitoring state while our infrastructure provider worked to fully resolve the underlying issue.\n\nBy 22:35 UTC, our provider reported full restoration of service. After validating stability in US-East, we completed rebalancing traffic by 23:48 UTC, and declared the incident resolved.\n\n### **Mitigation Steps and Recommended Future Preventative Measures**\u00a0\n\nWhile this incident was caused by an external infrastructure outage, we\u2019ve identified several opportunities to strengthen our internal readiness and response procedures.\n\nWe are consolidating and centralizing our regional failover procedures to ensure they are immediately accessible and complete for all production services. Any gaps in our process documentation for newer services will be addressed to ensure readiness before they are fully adopted into production. Additionally, we are reviewing and resolving issues with internal tooling, including inventory and DNS resolution problems, which made mitigation more difficult during the incident.\n\nThese improvements will ensure faster and more consistent responses to future infrastructure-level disruptions, and reduce potential impact on customer traffic across regions.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-10-20T09:19:51.504Z",
"resolved_inferred": false,
"started_at": "2025-10-20T07:32:41.000Z",
"state": "postmortem",
"title": "Elevated latencies and errors for multiple services in US-west and US East",
"updated_at": "2025-10-27T18:40:21.569Z",
"url": "https://stspg.io/2812t6gsfk7z"
},
{
"body": "### **Problem Description, Impact, and Resolution**\u00a0\n\nOn October 17, 2025 at 04:51 UTC, some customers may have experienced elevated latency and error rates with the Pub/Sub service in the IAD region \\(US-East\\). Our engineering teams began immediate investigation and identified a spike in errors related to a recent update to the Pub/Sub service.\n\nWe began formal incident response and initiated rollback of the service deployment shortly thereafter. The issue was fully resolved by 06:50 UTC, and rollback across all regions was completed by 08:00 UTC.\n\nThe issue occurred because a misconfiguration in the release caused incorrect behavior in the channel cleanup logic. Additionally, our alerting configuration did not include coverage for the synthetic test failures that would have surfaced this issue sooner, delaying detection.\n\n### **Mitigation Steps and Recommended Future Preventative Measures**\u00a0\n\nTo prevent a similar issue from occurring in the future, our engineering teams have written a simpler and more reliable replacement for the faulty logic. That code is currently undergoing rigorous testing before being reintroduced in a future release.\n\nWe are also addressing the lack of proper alerting that contributed to a delayed response. Synthetic tests have been reviewed, and appropriate alerting will be implemented to ensure similar regressions are detected earlier. In parallel, we are updating our development and testing processes to catch such issues before code reaches production. Lastly, we are conducting a refresher training on our incident response process to ensure faster execution and coordination in the future.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-10-17T08:06:42.282Z",
"resolved_inferred": false,
"started_at": "2025-10-17T06:42:58.000Z",
"state": "postmortem",
"title": "Potential for some missed messages for subscribers in IAD",
"updated_at": "2025-10-22T22:22:35.143Z",
"url": "https://stspg.io/sp381dwkw3yz"
},
{
"body": "### **Problem Description, Impact, and Resolution**\u00a0\n\nAt 08:40 UTC on October 6, 2025, we observed elevated error rates in the Events & Actions service in our EU-Central \\(FRA\\) region, which led to delays in processing publish-triggered events. Some customers may have experienced slower-than-expected execution of their event workflows during this time.\n\nWe identified a malformed payload that was causing backend consumers to fail when attempting to process the queue. We deployed an updated build with improved parsing logic, which cleared the blockage and restored normal service. The issue was fully resolved by 11:00 UTC on October 6, 2025.\n\nThis issue occurred because our event processing service did not correctly handle a malformed message format, which caused the processing queue to stall. Additionally, the alerting system in place was not configured to detect this failure mode promptly, delaying our response.\n\n### **Mitigation Steps and Recommended Future Preventative Measures**\u00a0\n\nTo reduce the risk of similar delays in the future, we are refining our alert thresholds and naming conventions to improve early detection and clarity during response. We are also reviewing validation logic to ensure malformed messages are consistently isolated before reaching backend queues.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-10-06T11:00:14.418Z",
"resolved_inferred": false,
"started_at": "2025-10-06T10:13:39.359Z",
"state": "postmortem",
"title": "Elevated Event & Action Error In FRA Region",
"updated_at": "2025-10-06T20:50:21.601Z",
"url": "https://stspg.io/8gpzcwlbqs91"
},
{
"body": "### **Problem Description, Impact, and Resolution**\u00a0\n\nAt 18:14 UTC on September 7, 2025 we observed increased error rates and latency for our Presence service in our San Jose, Virginia, and Tokyo regions. We increased capacity in those regions and the issue was resolved at 18:17 UTC. This issue was a recurrence of the issue [experienced on September 2, 2025](https://status.pubnub.com/incidents/1n8xk6w5y9lk), where a bug in one of our APIs allowed a request to execute an operation that exceeded assumed limits in extreme cases, causing out-of-memory conditions for the Presence service.\n\n### **Mitigation Steps and Recommended Future Preventative Measures**\u00a0\n\nIn the previous instance of this issue, we placed restrictions on the API in question; those changes were not restrictive enough, which allowed for this recurrence. We have corrected that oversight, as well as increased memory capacities in this area of our system as an additional safeguard.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-09-07T19:02:56.106Z",
"resolved_inferred": false,
"started_at": "2025-09-07T18:45:26.873Z",
"state": "postmortem",
"title": "Increased Error Rate and Latency for Presence",
"updated_at": "2025-09-12T23:30:21.453Z",
"url": "https://stspg.io/647gv2199kwm"
},
{
"body": "### **Problem Description, Impact, and Resolution**\u00a0\n\nAt 18:09 UTC on September 2, 2025 we observed increased error rates and latency for our Presence service in our San Jose, Virginia, and Tokyo regions. We increased capacity in those regions and the issue was resolved at 18:16 UTC. This issue occurred because a bug in one of our APIs allowed a request to execute an operation that exceeded assumed limits in extreme cases. In this case, a large number of such requests were executed that resulted in out-of-memory conditions for the Presence service.\n\n### **Mitigation Steps and Recommended Future Preventative Measures**\u00a0\n\nTo prevent a similar issue from occurring in the future we have enforced the intended limit on the API in question. We have also added additional testing and monitoring in this area.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-09-02T18:50:57.453Z",
"resolved_inferred": false,
"started_at": "2025-09-02T18:34:31.140Z",
"state": "postmortem",
"title": "Presence is experiencing elevated latencies and error rates",
"updated_at": "2025-09-05T22:40:21.488Z",
"url": "https://stspg.io/r89qvm24jzdv"
}
]
}