{
"vendor": "Pusher",
"slug": "pusher",
"platform": "statuspage",
"status_url": "https://status.pusher.com",
"last_checked": "2026-09-16T12:28:20Z",
"last_state": "ok",
"history_backfilled": true,
"first_watched": "2026-09-04T07:06:16Z",
"incidents": [
{
"body": "This incident has been resolved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-08-23T05:12:25.545Z",
"resolved_inferred": false,
"started_at": "2026-08-23T01:27:50.123Z",
"state": "resolved",
"title": "Elevated Errors on us2",
"updated_at": "2026-08-23T05:12:25.562Z",
"url": "https://stspg.io/zwqg1pjx8hfy"
},
{
"body": "This incident has been resolved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-08-22T04:47:53.810Z",
"resolved_inferred": false,
"started_at": "2026-08-22T03:17:10.408Z",
"state": "resolved",
"title": "Elevated errors on us2 cluster",
"updated_at": "2026-08-22T04:47:53.826Z",
"url": "https://stspg.io/7bnbcgkbhdkl"
},
{
"body": "This incident has been resolved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-04-06T02:55:49.359Z",
"resolved_inferred": false,
"started_at": "2026-04-06T01:52:05.304Z",
"state": "resolved",
"title": "Intermittent API errors in US2 cluster",
"updated_at": "2026-04-06T02:55:49.373Z",
"url": "https://stspg.io/80bmtqwc23fb"
},
{
"body": "This incident has been resolved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-03-31T12:44:41.147Z",
"resolved_inferred": false,
"started_at": "2026-03-31T12:33:28.286Z",
"state": "resolved",
"title": "Cluster SA1: Delayed Channels webhook delivery on SA1 cluster",
"updated_at": "2026-03-31T12:44:41.163Z",
"url": "https://stspg.io/mh5dx011mv9x"
},
{
"body": "This incident has been resolved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-02-15T22:14:20.919Z",
"resolved_inferred": false,
"started_at": "2026-02-15T20:33:40.188Z",
"state": "resolved",
"title": "Increased error rate on the US2 cluster",
"updated_at": "2026-02-15T22:14:20.940Z",
"url": "https://stspg.io/0rv44htxtw4n"
},
{
"body": "This incident has been resolved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-01-29T10:58:45.878Z",
"resolved_inferred": false,
"started_at": "2026-01-29T10:38:02.971Z",
"state": "resolved",
"title": "Errors increase for API in cluster us3",
"updated_at": "2026-01-29T10:58:45.894Z",
"url": "https://stspg.io/89k2khjnhry9"
},
{
"body": "This incident has been resolved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-01-26T22:43:12.072Z",
"resolved_inferred": false,
"started_at": "2026-01-26T22:29:04.676Z",
"state": "resolved",
"title": "Increased error rate on the US2 cluster impacting presence events",
"updated_at": "2026-01-26T22:43:12.087Z",
"url": "https://stspg.io/b3420jvz2lk3"
},
{
"body": "All affected accounts have been restored to their correct subscription levels. We sent an email yesterday to all impacted customers. If you notice any issues with your subscription, please contact us.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-01-23T11:58:11.160Z",
"resolved_inferred": false,
"started_at": "2026-01-22T16:38:00.329Z",
"state": "resolved",
"title": "Incorrect subscription changes affecting some accounts",
"updated_at": "2026-01-23T11:58:11.186Z",
"url": "https://stspg.io/12sq01t39cdc"
},
{
"body": "This incident has been resolved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-01-16T12:41:18.954Z",
"resolved_inferred": false,
"started_at": "2026-01-16T08:48:53.259Z",
"state": "resolved",
"title": "Intermittent errors with message delivery on cluster us2",
"updated_at": "2026-01-16T12:41:18.969Z",
"url": "https://stspg.io/9cw9jg93w3wt"
},
{
"body": "This incident has been resolved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-01-16T01:37:57.048Z",
"resolved_inferred": false,
"started_at": "2026-01-16T01:07:12.399Z",
"state": "resolved",
"title": "Intermittent errors with message delivery",
"updated_at": "2026-01-16T01:37:57.069Z",
"url": "https://stspg.io/hvmt3rcbnyfb"
},
{
"body": "This incident has been resolved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-10-27T13:34:34.536Z",
"resolved_inferred": false,
"started_at": "2025-10-27T11:50:02.599Z",
"state": "resolved",
"title": "AP1 cluster - Socket connection failures",
"updated_at": "2025-10-27T13:34:34.551Z",
"url": "https://stspg.io/kcbvc21rrbh3"
},
{
"body": "## **Root Cause Analysis: Elevated API Errors and Outage in AP4 Cluster**\n\n**Incident Date:** October 20, 2025\n\n**Status:** Resolved\n\n### **Summary**\n\nBetween **October 20 and October 21, 2025**, customers using the **AP4 cluster** experienced elevated API errors, latency, and message publishing failures. The issue primarily affected the **Channels API**, preventing customers from publishing new messages and leading to degraded real-time functionality for end-users.\n\nDuring system recovery and while implementing mitigations from a previous incident on Oct 18th, a **misconfigured Redis container** in the AP4 cluster failed to start correctly, preventing caching operations needed for API requests. This misconfiguration went undetected by proactive monitoring, delaying full recovery until October 21 at 17:43 UTC.\n\n### **Impact**\n\nThroughout the incident period, customers in the **AP4 cluster** experienced:\n\n* **High API error rates** when attempting to publish messages through the Channels API\n* **Failed or delayed message delivery** for connected clients\n* **Temporary downtime** for end-customer applications relying on real-time messages\n\nOther clusters remained operational, though some minor latency was observed in isolated regions due to dependencies on shared services.\n\n### **Root Cause**\n\nThis incident resulted from **a chain of events** involving both external and internal factors:\n\n1. \\*\\*Major AWS Outage \\(October 20\\)\\*\\*A large-scale **AWS outage in the US-East region** disrupted multiple dependent systems, impacting several Pusher clusters.\n2. **Misconfigured Redis Container \\(October 21\\)** As systems in the AP4 cluster attempted to scale during recovery, one of the backend **Redis cache containers** failed to start due to a **misconfigured environment variable**. This prevented Redis from initializing properly, resulting in API operations failing or timing out.\n3. **Monitoring Gap** Existing monitoring did not capture the **Redis startup failure** because the specific failure mode occurred after initialization checks had passed. This delayed internal detection until API error rates increased and customer impact was observed.\n4. **Delayed Customer Communication** Initial updates to customers were delayed while the team triaged the issue and verified the failure pattern, prolonging the time before external notification.\n\n### **Detection and Response**\n\nThe issue was detected through a combination of **monitoring alerts** showing elevated error rates and **customer reports** of publishing failures.\n\n**Timeline of Events:**\n\n* **October 20** \u2013 AWS outage began, affecting multiple Pusher clusters leading to increased delays and errors. **October 20, evening UTC** \u2013 Pusher clusters began recovery as AWS services were restored.\n* **October 21, 15:06 UTC** \u2013 Internal monitoring detected elevated API errors in AP4; engineers began investigation.\u00a0 Incident was unrelated to prior AWS outage.\n* **October 21, 15:09 UTC** \u2013 Root cause identified as a failed Redis caching container.\n* **October 21, 15:27 UTC** \u2013 Restoration of Redis connections underway.\n* **October 21, 15:29 UTC** \u2013 Fix implemented; cluster began gradual recovery.\n* **October 21, 17:39 UTC** \u2013 Full stabilization confirmed across AP4 nodes.\n* **October 21, 17:43 UTC** \u2013 Incident marked resolved after sustained recovery.\n\n### **Resolution**\n\nTo restore full functionality, the engineering team:\n\n* Corrected the **Redis container configuration** preventing startup\n* Restarted and validated cache services across all AP4 nodes\n* Confirmed API endpoints were fully operational and message publishing resumed\n* Monitored latency and error metrics to confirm sustained stability\n\n### **Preventative Actions**\n\nTo reduce recurrence risk and improve detection and response, Pusher is implementing the following:\n\n* **Enhanced Redis Monitoring:** Extending monitoring coverage to detect Redis startup and post-init failures.\n* **Customer Communication Enhancements:** Improving internal escalation and communication processes to ensure faster external updates.\n\n### **Next Steps and Commitment**\n\nWe recognize the importance of reliable API performance for our customers. Our teams are conducting a full review of caching dependencies and configuration management across all clusters to prevent similar incidents.\n\nWe sincerely apologize for the disruption caused by this event and appreciate your patience as we worked through a complex multi-day recovery scenario. Pusher remains committed to transparency, reliability, and continuous improvement in service resilience.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-10-21T17:43:25.752Z",
"resolved_inferred": false,
"started_at": "2025-10-21T15:06:39.819Z",
"state": "postmortem",
"title": "Elevated API Errors in AP4 Cluster",
"updated_at": "2025-10-23T14:14:30.222Z",
"url": "https://stspg.io/zcbl5knddsk2"
},
{
"body": "## Root Cause Analysis: Redis Cluster Startup Failures \u2013 US3 Cluster\n\n**Incident Date:** October 20, 2025    \n**Duration:** 18:36 UTC \u2013 20:41 UTC    \n**Status:** Resolved\n\n## **Summary**\n\nOn October 20, 2025, customers experienced a significant outage affecting the US3 cluster.\n\nThe disruption began with increased latency and 503 errors before escalating into full service downtime as Redis clusters in the US3 region failed to start successfully during an infrastructure update.\n\nService was fully restored at 20:41 UTC after engineers identified and resolved the underlying startup issue, confirming stability across all Redis clusters.\n\n## **Impact**\n\nBetween **18:36 UTC and 20:41 UTC**, customers experienced:\n\n* Major outage in the **US3** cluster, impacting message delivery and connection reliability\n* Elevated error rates and timeouts across dependent APIs\n* Temporary need for customers to switch to alternate clusters to maintain service continuity\n\nNo customer data was lost. However, applications relying solely on the affected cluster experienced full downtime for a large portion of the incident.\n\n## **Root Cause**\n\nThe outage was caused by a failure of Redis clusters to start correctly following an infrastructure update due to a missing configuration flag upon startup.\n\nAn unexpected upgrade to the Docker runtime running our Redis cluster introduced a breaking change that prevented container startup for certain Redis deployments. When replacement Redis instances in the US3 redis-main cluster were launched, they failed initialization checks and repeatedly restarted, rendering the cluster unavailable.\n\nThe incompatibility remained undetected until a routine node replacement in the US3 cluster introduced new Redis instances to the cluster.\n\n## **Detection and Response**\n\nMonitoring systems first detected increased error rates and latency at **18:36 UTC**, followed by a rise in 503 responses from the affected APIs.\n\n**Timeline of events:**\n\n* **18:36 UTC** \u2013 Increased latency and 503 errors observed in US3 cluster\n* **19:01 UTC** \u2013 Engineering began investigation into Redis startup failures\n* **19:04 UTC** \u2013 Incident declared a major outage; mitigation efforts initiated\n* **19:59 UTC** \u2013 Root cause identified and configuration fix applied\n* **20:20 UTC** \u2013 Services operational; monitoring for recovery stability\n* **20:41 UTC** \u2013 All Redis nodes confirmed healthy; incident resolved\n\n## **Resolution**\n\nThe engineering team:\n\n* Implemented a temporary fix to the Redis environments ensure Redis instances could initialize successfully\n* Blocked further automated replacements in other clusters until validated\n* Verified recovery and stability across all Redis clusters\n\nAfter deployment of the fix, Redis instances in the US3 redis-main cluster started correctly and full service was restored.\n\n## **Preventative Actions**\n\nTo prevent recurrence, the team has:\n\n* Rolled out a permanent fix across all Redis clusters used in all regions\n* Planned a long-term remediation to modernize Redis image packaging for compatibility with current and future Docker releases\n\n## **Next Steps and Commitment**\n\nWe are conducting a broader review of infrastructure upgrade processes to better detect runtime incompatibilities before they impact production.\n\nWe apologize for the disruption this incident caused. Ensuring reliability and transparency remains our highest priority, and we continue to strengthen our processes to maintain consistent, predictable service for all Pusher customers.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-10-20T20:41:37.266Z",
"resolved_inferred": false,
"started_at": "2025-10-20T18:36:06.632Z",
"state": "postmortem",
"title": "US3 cluster - major outage",
"updated_at": "2025-11-03T20:05:40.399Z",
"url": "https://stspg.io/dcfktz1mqkps"
},
{
"body": "This incident has been resolved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-10-21T07:14:52.513Z",
"resolved_inferred": false,
"started_at": "2025-10-20T17:26:18.885Z",
"state": "resolved",
"title": "MT1 cluster - increased latency",
"updated_at": "2025-10-21T07:14:52.529Z",
"url": "https://stspg.io/vth06sd5l50v"
},
{
"body": "This incident has been resolved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-10-21T07:13:16.479Z",
"resolved_inferred": false,
"started_at": "2025-10-20T17:21:26.390Z",
"state": "resolved",
"title": "Pusher Beams degradation",
"updated_at": "2025-10-21T07:13:16.496Z",
"url": "https://stspg.io/732gkz4w3x45"
},
{
"body": "Our cloud provider is reporting recovery across most affected services. On our side, we haven\u2019t noticed any issues in the past hour, and all services are operating normally. We continue to monitor the situation closely and will provide updates if anything changes.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-10-20T10:18:27.948Z",
"resolved_inferred": false,
"started_at": "2025-10-20T08:19:49.127Z",
"state": "resolved",
"title": "Beams - webhook delivery issues",
"updated_at": "2025-10-20T10:18:27.967Z",
"url": "https://stspg.io/nt3cyt19x0cy"
},
{
"body": "This incident has been resolved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-10-20T12:57:01.049Z",
"resolved_inferred": false,
"started_at": "2025-10-20T08:10:37.068Z",
"state": "resolved",
"title": "MT1 cluster - increased latency and webhook delivery failures",
"updated_at": "2025-10-20T12:57:01.069Z",
"url": "https://stspg.io/8ww72kngbflv"
},
{
"body": "## **Root Cause Analysis: Increased Latency and Message Delivery Failures \u2013 US2 Cluster**\n\n**Incident Date:** October 19, 2025\n\n**Duration:** 00:44 UTC \u2013 08:17 UTC\n\n**Status:** Resolved\n\n### **Summary**\n\nOn October 19, 2025, customers using Pusher experienced increased latency and message delivery failures. These issues primarily affected the US2 cluster, with intermittent impact also observed in the MT1 and US3 clusters.\u00a0 The incident resulted in delayed or undelivered messages for many applications.\n\nLatency stabilized at 08:17 UTC after mitigation actions were completed.\n\n### **Impact**\n\nBetween 00:44 UTC and 08:17 UTC, multiple customers experienced:\n\n* **Delayed or failed message delivery** across affected clusters\n* **Degraded performance** in connection establishment and publishing\n\nThe most significant and prolonged impact occurred in the **US2 cluster**, while **MT1** and **US3** clusters saw elevated latency for a shorter period before stabilizing.\n\n### **Root Cause**\n\nThe primary cause of the incident was IP address saturation within the subnet assigned to the public Pusher clusters.\n\nWhen traffic levels increased, the **US2 cluster** was unable to scale out further because the available IP addresses in its subnet were fully utilized. This IP scaling limitation prevented the creation of additional instances needed to handle the load.\n\nSecondary factors included temporary **network saturation** and **capacity limits** at our cloud provider, which amplified the latency in the early stages of the incident.\n\n### **Detection and Response**\n\nThe issue was first detected through a combination of **customer reports** and **internal monitoring alerts** showing elevated response times and connection errors.\n\nThe timeline of actions was as follows:\n\n* **00:44 UTC** \u2013 Monitoring alerted the team to increased latency across MT1, US2, and US3 clusters.\n* **03:11 UTC** \u2013 Engineers identified subnet capacity as a contributing factor; mitigations began.\n* **04:03 UTC** \u2013 All clusters were scaled out to distribute traffic; latency began to improve in MT1 and US3 clusters.\n* **08:17 UTC** \u2013 Manual intervention allowed US2 to scale successfully, restoring normal latency.\n\n### **Resolution**\n\nTo restore service, the engineering team:\n\n* Scaled out **MT1** and **US3** clusters to handle increased traffic loads\n* Monitored all clusters to confirm sustained stability\n\nOnce additional capacity was provisioned and high loads normalized, latency levels returned to normal and remained stable.\n\n### **Preventative Actions**\n\nTo prevent recurrence, Pusher has initiated the following actions:\n\n* **Rate Limits:** We will re-evaluate how rate limits are implemented to better mitigate content from neighboring customers on the shared clusters.\n* **Subnet Expansion:** Re-evaluating and increasing the size of subnets assigned to shared clusters to ensure sufficient IP availability for future scaling events.\n* **Load Balancer Enhancements:** We will implement load balance sharding in order to better distribute connections.\n* **Capacity Planning Improvements:** Enhancing internal monitoring and alerting for subnet and IP utilization thresholds.\n\n### **Next Steps and Commitment**\n\nWe recognize that message latency and delivery reliability are critical to our customers\u2019 applications. Our team is continuing a full review of cluster capacity management and provider configuration to improve resilience under high traffic conditions.\n\nWe apologize for the disruption this incident caused and appreciate your patience while we worked to resolve it. Ensuring reliability and transparency remains our highest priority.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-10-19T08:17:11.384Z",
"resolved_inferred": false,
"started_at": "2025-10-19T00:44:46.098Z",
"state": "postmortem",
"title": "Increased latency - US2 cluster",
"updated_at": "2025-10-23T14:17:37.091Z",
"url": "https://stspg.io/cqw0hdd2dqmy"
},
{
"body": "This incident has been resolved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-09-21T07:54:53.140Z",
"resolved_inferred": false,
"started_at": "2025-09-19T21:15:37.686Z",
"state": "resolved",
"title": "Higher latency on our AP1 cluster for some customers",
"updated_at": "2025-09-21T07:54:53.161Z",
"url": "https://stspg.io/7p71ldc2pzzn"
}
]
}