{
"vendor": "Courier",
"slug": "courier",
"platform": "statuspage",
"status_url": "https://status.courier.com",
"last_checked": "2026-09-16T12:28:20Z",
"last_state": "ok",
"history_backfilled": true,
"first_watched": "2026-09-04T07:06:16Z",
"incidents": [
{
"body": "Send pipeline has resolved and enqueued messages are now sending properly.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-06-18T15:09:09.632-07:00",
"resolved_inferred": false,
"started_at": "2026-06-18T13:11:52.004-07:00",
"state": "resolved",
"title": "Send Pipeline Latency",
"updated_at": "2026-06-18T15:09:09.648-07:00",
"url": "https://stspg.io/01f7ydq6xxlw"
},
{
"body": "Please continue to monitor Slack's statuspage for any updates related to message deliverability: https://slack-status.com/",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-05-27T16:16:34.216-07:00",
"resolved_inferred": false,
"started_at": "2026-05-27T16:06:26.000-07:00",
"state": "resolved",
"title": "Slack Deliverability Issues",
"updated_at": "2026-05-27T16:16:34.233-07:00",
"url": "https://stspg.io/8vdrrzbyhrym"
},
{
"body": "# RFO: February 2, 2026 \u2014 Service Interruption\n\n## Executive Summary\n\nOn February 2, 2026, a deployment introduced a configuration change that referenced a file not present in our production build artifacts, causing backend services to become unavailable. API endpoints, automations, and tracking functionality were impacted for 2 hours and 17 minutes in the US region and 2 hours and 42 minutes in the Ireland region. Recovery was prolonged by a concurrent outage at GitHub Actions, our CI/CD provider, which prevented our standard automated rollback. Our team identified the root cause, executed a manual rollback independent of the affected provider, and restored full service across all regions. We have defined targeted action items to prevent recurrence.\n\n## Incident Overview\n\n* **Affected services:** API endpoints \\(including message sending\\), automations, webhooks, tracking links, and authentication\n* **Impact:** Requests to backend services returned errors for the duration of the incident. Users were unable to send messages, trigger automations, or access tracking data. No data was lost \u2014 requests were rejected before ingestion, so no messages were partially processed or left in an inconsistent state.\n* **Detection:** Our monitoring systems flagged elevated error rates within minutes of the issue beginning.\n* **Contributing factor:** GitHub Actions, our CI/CD provider, experienced a complete outage from 10:35 to 16:30 PST. This overlapped with our incident window and prevented our standard automated rollback from executing, extending the time to resolution.\n\n## Timeline of Events\n\nAll times in PST.\n\n| **Time** | **Event** |\n| --- | --- |\n| 10:59 | Deployment of latest release initiated through standard CI/CD pipeline |\n| 11:44 | Deployment completed and went live; services immediately began experiencing errors due to a missing configuration dependency |\n| 11:49 | GitHub Actions, our CI/CD provider, experienced a complete outage, preventing standard rollback procedures |\n| 11:52 | Monitoring alerts triggered; engineering team engaged |\n| 12:00 | Incident declared; rollback initiated; engineering team assembled |\n| 12:07 | Status page updated \u2014 issue identified and rollback in progress. Rollback ends up being blocked by GitHub actions outage. |\n| 12:39 | Team pivoted to an alternative manual rollback approach. This required testing and validating the new approach. |\n| 13:50 | Team executed the alternative manual rollback independent of GitHub Actions after they were fully satisfied that the new approach was safe.\u00a0 |\n| 14:01 | Manual rollback completed in US region; services confirmed operational |\n| 14:26 | Ireland region deployment completed; services confirmed operational |\n| 15:13 | All services verified stable across all regions; incident resolved |\n\n## Root Cause Analysis\n\nThe disruption was traced to a configuration change included in the latest release. The change introduced a startup dependency on a utility file that was intended to be bundled with the deployment package. However, the file was not included in the production build artifacts. When backend services attempted to initialize, they were unable to locate the required file and could not start, resulting in all incoming requests being rejected.\n\nThis discrepancy was not caught prior to production deployment because the file was present and functioning correctly in the development environment. The difference in how build artifacts are assembled between development and production environments meant the issue only manifested in production.\n\n## Mitigation and Resolution\n\n1. Upon identifying the root cause, the team immediately initiated a rollback to the prior known-good release through our standard CI/CD pipeline.\n2. A complete outage at GitHub Actions, our CI/CD provider, prevented the automated rollback from completing. The team identified this external dependency and pivoted to an alternative approach.\n3. The team executed a manual rollback by retrieving prior deployment artifacts from our backup storage and deploying them directly, bypassing the affected CI/CD pipeline entirely.\n4. Services were restored region by region, with the US region confirmed operational at 14:01 PST and the Ireland region at 14:26 PST.\n5. Extended monitoring was conducted across all services before declaring the incident fully resolved at 15:13 PST after we were satisfied everything had been stable for over 45 minutes.\u00a0\n\n## Action Items\n\n| **#** | **Action Item** | **Owner** | **Priority** | **Status** |\n| --- | --- | --- | --- | --- |\n| 1 | Require successful staging deployment and smoke tests before any production deployment | Engineering | P1 | In Progress |\n| 2 | Improve the reliability of automated smoke tests | Engineering | P1 | In Progress |\n| 3 | Add build-time validation to confirm all referenced startup dependencies are present in deployment packages | Engineering | P2 | Open |\n| 4 | Adopt an expedited rollback process as the standard emergency procedure, independent of CI/CD provider availability, reducing recovery time by approximately 35 minutes | Engineering | P2 | In Progress |",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "critical",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-02-02T15:13:46.580-08:00",
"resolved_inferred": false,
"started_at": "2026-02-02T12:07:08.817-08:00",
"state": "postmortem",
"title": "Courier Multi Service Outage",
"updated_at": "2026-02-19T09:37:35.357-08:00",
"url": "https://stspg.io/yr3h56fl3c1m"
},
{
"body": "The Courier team identified an issue affecting observability metrics between 21:38 UTC and 22:13 UTC, during which metrics briefly experienced an outage. A fix was released and metrics are stabilized through observability channels. Metrics received during the outage window will not be accounted for in observability dashboards.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-01-29T02:00:00.000-08:00",
"resolved_inferred": false,
"started_at": "2026-01-29T02:00:00.000-08:00",
"state": "resolved",
"title": "Observability Degradation",
"updated_at": "2026-01-29T15:13:40.011-08:00",
"url": "https://stspg.io/77yfmbx8sr7h"
},
{
"body": "Fix is out and application is stable",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-01-14T16:56:21.489-08:00",
"resolved_inferred": false,
"started_at": "2026-01-14T13:47:10.871-08:00",
"state": "resolved",
"title": "Courier Web App Performance",
"updated_at": "2026-01-14T16:56:21.504-08:00",
"url": "https://stspg.io/y64bqb8vzm25"
},
{
"body": "The incident has been resolved",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "critical",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-10-20T03:11:54.758-07:00",
"resolved_inferred": false,
"started_at": "2025-10-20T02:19:57.000-07:00",
"state": "resolved",
"title": "All services impacted",
"updated_at": "2025-10-20T03:12:47.604-07:00",
"url": "https://stspg.io/dbl8v74y9xns"
},
{
"body": "Release is in production and Microsoft Teams messages are passing successfully.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-10-16T23:01:48.017-07:00",
"resolved_inferred": false,
"started_at": "2025-10-16T22:33:05.648-07:00",
"state": "resolved",
"title": "Microsoft Teams Tenant Id Errors",
"updated_at": "2025-10-16T23:01:48.038-07:00",
"url": "https://stspg.io/9n639vr2wxj4"
},
{
"body": "Courier experienced a slowdown in message delivery with Mailgun providers around 2AM PST. Our monitoring systems caught the spike in the bottleneck, which eventually stabilized. Our infrastructure team is actively monitoring.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-10-13T09:19:58.000-07:00",
"resolved_inferred": false,
"started_at": "2025-10-13T09:19:58.000-07:00",
"state": "resolved",
"title": "Mailgun Message Delivery Latency",
"updated_at": "2025-10-13T12:50:20.930-07:00",
"url": "https://stspg.io/b5ly29z5s7sd"
},
{
"body": "Courier's delivery platform experienced significant latency around 1am PST. Our platform team managed to resolve the issue and is actively monitoring for any potential issues.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-10-08T02:00:00.000-07:00",
"resolved_inferred": false,
"started_at": "2025-10-08T02:00:00.000-07:00",
"state": "resolved",
"title": "Message Delivery Latency",
"updated_at": "2025-10-08T08:16:43.235-07:00",
"url": "https://stspg.io/2lfmxnrzn1c3"
}
]
}