{
"vendor": "Coveralls",
"slug": "coveralls",
"platform": "statuspage",
"status_url": "https://status.coveralls.io",
"last_checked": "2026-09-16T12:28:20Z",
"last_state": "ok",
"history_backfilled": true,
"first_watched": "2026-09-04T07:06:16Z",
"incidents": [
{
"body": "This incident has been resolved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-08-11T14:09:44.458-07:00",
"resolved_inferred": false,
"started_at": "2026-08-11T12:02:12.967-07:00",
"state": "resolved",
"title": "\"Website under heavy load\" warning",
"updated_at": "2026-08-11T14:09:44.475-07:00",
"url": "https://stspg.io/nd9v2n7g91wc"
},
{
"body": "This issue was resolved over the weekend. We will continue to monitor for elevated latency in the modified queues.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-20T10:32:17.056-07:00",
"resolved_inferred": false,
"started_at": "2026-07-15T10:28:59.136-07:00",
"state": "resolved",
"title": "Increased latency for large repos",
"updated_at": "2026-07-20T10:32:17.078-07:00",
"url": "https://stspg.io/fddc34gx1s03"
},
{
"body": "This incident has been resolved but we believe it triggered a worsened incident overnight Sun night/Mon morning (US PDT) which has just been resolved. To be confirmed by full RCA, we believe a deluge of large repo uploads tied up individual web servers that handle frontline requests. Each web server is able to recover on its own, and did, but as volume increased all servers were eventually affected, only allowing short windows where requests could get through\u2014rejecting most requests with 504 errors.\n\nWe'll post a post-mortem when we understand more about what happened and how to prevent it going forward.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-13T07:56:35.290-07:00",
"resolved_inferred": false,
"started_at": "2026-07-10T08:00:47.000-07:00",
"state": "resolved",
"title": "Increased latency for large repos",
"updated_at": "2026-07-13T07:56:35.309-07:00",
"url": "https://stspg.io/tyhgk4l60pfp"
},
{
"body": "Increased latency for large repos has been resolved for the general public. We will be clearing three outlier repos overnight, which should be fully cleared by tomorrow AM.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-07T21:39:37.170-07:00",
"resolved_inferred": false,
"started_at": "2026-07-07T09:43:22.969-07:00",
"state": "resolved",
"title": "Increased latency for large repos",
"updated_at": "2026-07-07T21:39:37.187-07:00",
"url": "https://stspg.io/rsxh917p1sk6"
},
{
"body": "We are closing this incident. Our main background processing queue for larger repos has had no backups in 48 hrs. We will continue monitoring for recurrence.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-05-28T08:53:52.401-07:00",
"resolved_inferred": false,
"started_at": "2026-05-18T09:41:15.110-07:00",
"state": "resolved",
"title": "Increased latency for large projects",
"updated_at": "2026-05-28T08:53:52.415-07:00",
"url": "https://stspg.io/84xn0rdzf008"
},
{
"body": "This incident has been resolved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-05-10T08:11:41.374-07:00",
"resolved_inferred": false,
"started_at": "2026-05-10T07:44:16.282-07:00",
"state": "resolved",
"title": "Unschedule maintenance",
"updated_at": "2026-05-10T08:11:41.388-07:00",
"url": "https://stspg.io/gspzxkkdmphc"
},
{
"body": "# Suspended for Not Paying\u2014While Paying\n\n_Latest update: Tuesday, Mar 24_\n\n[This post-mortem is also available as a PDF](https://s3.amazonaws.com/assets.coveralls.io/statuspage/incidents/20260224/postmortem/Coveralls-Post-Mortem-Service-Outage-Feb-2026.pdf).\n\n## Overview\n\nOn February 24, our hosting provider suspended our account. Our service was down for 68 hours. The stated reason was past-due charges. When we reached our account manager on the day of suspension, he reviewed our account and told us what he saw: an \"account restriction\" that, in his words, \"shouldn't have happened.\"\n\nBizarrely, his review turned up an internal note claiming we'd made _no payments in six months_. That was demonstrably false and we provided immediate evidence to prove it.\n\nWe did have an outstanding balance\u2014the product of a billing dispute that took nearly a full year to reach a decision in our favor. But finalizing what we owed required tracking down two credits we couldn't locate. We asked AR one specific question: _where did those two credits go?_ We asked it _six times_. For some reason we still don't understand, they simply stopped responding. We escalated to our account manager, who promised to intervene. Then we heard nothing from anyone. Until we were suspended.\n\nDespite providing evidence of our payments, open billing cases, and relentless communication, we waited three days for our service to be restored. When someone finally acted, we were back online within minutes. Three days, and then minutes. Our provider has not explained that gap, or responded to the formal request for answers we submitted more than two weeks ago.\n\nWithout those answers, what follows is our account. We're sharing everything and letting you judge for yourself.\n\n_If what you need most right now is assurance that this won't happen again, and to know what we're doing for affected customers, you can jump ahead to **What We're Doing** and **Making It Right**._\n\nBut we'd ask you to read through. On the surface, this looks like a company that didn't pay its bills and got suspended. The full picture tells a different story, and the evidence speaks for itself.\n\n## About Coveralls\n\nCoveralls has been providing code coverage analytics for 15 years. We're bootstrapped: no outside funding. Over 225,000 developers use [Coveralls.io](http://Coveralls.io) every day, and roughly 90% of them are open-source contributors who use the service for free. \\[1\\]\n\nWe have never experienced an outage like this in our 15-year history.\n\n## The Incident\n\n### What We Were Told About Suspension\n\nMore than once, over the past year, we asked our account manager directly whether we were at risk of suspension.\n\nEach time, the answer was no. He told us suspension was designed to be rare, that it was reserved for accounts that had gone _delinquent_\u2014stopped _paying_, and stopped _communicating_\u2014and that it required his sign-off, _and_ his manager's.\n\nHe told us he received a regular list of at-risk accounts among those he managed, so he could review them and give feedback before any action was taken. He told us we had _never_ appeared on that list.\n\nHe said our provider _knew_ we were paying, and _knew_ we were actively communicating, with AR and with him. Based on everything he told us, suspension wasn't just unlikely, it wasn't supposed to be possible. At least not without warning.\n\n### Timeline\n\n**Tuesday, Feb 24 \u2014 Day 1**\n\n1:45 PM PST: We suddenly lose access to our hosting provider's console. The service is still up, so we assume it's a technical issue. Multiple team members try their credentials. None work.\n\n2:42 PM: We receive our first support request saying, \"posting coverage data stopped working\", but we don't see it until 2:51 PM due to the next event.\n\n2:45 PM: A team member gets past a login screen and sees a message: \"Your account was suspended because of past due charges.\" We don't believe it. We reach out to our account manager by email, then text, then phone. He confirms an \"account restriction\" that, based on his review of our account, \"shouldn't have happened\", and shares that an internal note on our account says, \"_Customer has made no payments for 6 months_.\"\n\n2:51 PM: We see the first support request.\n\n2:52 PM: We open an incident on our status page. At this point, we are still expecting this to be reversed quickly and do not yet disclose the suspension publicly. In hindsight, we should have.\n\n~3:00 PM: While we're on the phone with our account manager, one of us discovers the official suspension notification in our inbox, timestamped 2:03 PM, about 15 minutes after we were first locked out. Our AM asks us to create a new support ticket and promises to escalate. Two hours fly by while we feverishly collect and send evidence of improper suspension. We expect our AM to get this reversed as soon as he gets through to\u2014_someone_.\n\n5:36 PM: No definitive update. Our AM is \"still working on it.\" He tells us the AR team involved is on Eastern time and may be gone for the day. We post that the issue has been identified and a fix is being implemented. We believe this. We are wrong.\n\n7:28 PM: No resolution. Our account manager tells us he continues to escalate but that under the current circumstances, he himself has \"limited access\" to our account. Still\u2014_should_ be fixed by morning. We update our incident, extend the outage window, and recommend fail-on-error: false for CI workflows.\n\n~9:30 PM: Still no resolution. We check status throughout the night.\n\n**Wednesday, Feb 25 \u2014 Day 2**\n\n6:45 AM: We text our account manager, surprised service has not been restored.\n\n9:48 AM: First public disclosure \u2014 outage originated at our infrastructure provider. We are unable to access our account.\n\n5:56 PM: After a full day working with our provider's account team, we learn more about the restriction. Not a technical failure\u2014infrastructure and data intact. Their team indicates they are working to reverse it, and we expect resolution by morning.\n\n**Thursday, Feb 26 \u2014 Day 3**\n\n6:53 AM: Still waiting. \"We thought this would be resolved by this morning.\"\n\n5:13 PM: After a full second day of trying to get answers as to why this is happening, and why it hasn't been resolved yet, the picture remains unclear. No ETA. \"What appears to have been an automated action, one that no one on their side has taken ownership of causing or fixing.\" We commit to evaluating backup infrastructure and promise a full post-mortem.\n\n**Friday, Feb 27 \u2014 Day 4**\n\n7:41 AM: \"As we pass milestones in this outage that we never imagined we would see \\(and haven't in 15 years of operation\\)...\" We announce that if not restored by EOD, we will stand up new infrastructure over the weekend.\n\n11:02 AM: Systems begin coming back online. We have not been told what changed.\n\n11:11 AM: Service partially restored.\n\n12:23 PM: After one hour of monitoring, cautiously optimistic that service is fully restored in all regions.\n\n**2:30 PM: After two more hours of monitoring, the incident is officially marked resolved. Total outage: approximately three days \\(68 hours\\) across 4 calendar days.**\n\n\u200c\n\n## What We Think Happened\n\nHere's what we found when we started looking for answers:\n\n### The Note That Wasn't True\n\nThe first thing we learned about why our account had been suspended was what our account manager told us: there was an internal note from AR that read:\n\n> _Customer has made no payments for 6 months._\n\nThis was demonstrably false. Here is our payment history for that exact period\u2014September 1, 2025 through February 24, 2026 \\(the date of our suspension\\):\n\n![](https://s3.amazonaws.com/assets.coveralls.io/statuspage/incidents/20260224/postmortem/image2.png)\n\nWe shared this evidence with our provider immediately, backed by bank statements:\n\n![](https://s3.amazonaws.com/assets.coveralls.io/statuspage/incidents/20260224/postmortem/image1.png)\n\n\u200c\n\nBut it raises the question: if we were paying, why would our provider think we _weren't_?\n\n### Why The Note Existed\n\nWhen we look at our billing console we can see why someone might _think_ we hadn't paid: **it shows none of our payments**.\n\nInvoices we know to have been fully paid show none of our payments against them.\n\nMost notably, the **Transactions Tab** of the **Payments** section of our billing console contains _**no payment records**_ **for last 6 months**\u2014resonating with that internal note:\n\n\u200c\n\n\u200c\n\n![](https://s3.amazonaws.com/assets.coveralls.io/statuspage/incidents/20260224/postmortem/image4.png)\n\u200c\n\nIf an automated system relied on this same data to determine whether a customer had made payments, and that data showed six months of silence, it _might_ trigger exactly the kind of suspension we experienced: an automated action with no advance warning, with no one on their side able to explain it or willing to take ownership of it.\n\nThis is our strongest piece of evidence that the suspension was triggered by bad data, and not our actual account status.\n\n\u2014\n\nTo put this in context: over the six months during which the internal note claimed we'd paid _nothing_, our provider billed us $127,186. Our verified payments for the same period totaled $120,312\u2014a gap of roughly $6,900, which is less than a third of the amount we were actively disputing in our second open billing case.\n\n![](https://s3.amazonaws.com/assets.coveralls.io/statuspage/incidents/20260224/postmortem/image3.png)\n\n\n![](https://s3.amazonaws.com/assets.coveralls.io/statuspage/incidents/20260224/postmortem/image5.png)\n\n\\(By March 5, our payments even exceeded our total invoiced charges for the period: we were not delinquent.\\)\n\nOur payments were in line with our usage. We weren't _up-to-date_\u2014but the portion of the balance we still carried traced directly back to a single incident that triggered a billing dispute, which took ten months to resolve and still hadn't fully landed by the time we were suspended.\n\n\u2014\n\nBut the note, and the billing console, only explain so much. They might explain how the suspension was _triggered_. They don't explain why it took _three days_ to reverse, even after providing evidence of our payments and open billing cases. If there's an answer for that, it lives in the rest of the story. We've looked. We don't see a justification for either. Draw your own conclusions.\n\n### How The Balance Happened\n\nIn 2025, we faced a critical infrastructure challenge. Our production PostgreSQL database, approaching a 65TB ceiling due to failing maintenance operations, also needed a major version upgrade before end-of-life. We started the upgrade a month early. It took four-and-a-half months to complete.\n\nDuring three-and-a-half of those months, our provider levied Extended Support Fees on us that roughly _quadrupled_ our database hosting costs for that period. \\[2\\] We asked our provider to pause or waive them\u2014the fees were simply beyond what we could absorb. They told us that wasn't possible, but that we could open a billing dispute and make our case. We did that, and our provider eventually approved credits for the full amount. But that process took ten months. \\[3\\]\n\nDuring those ten months, those fees accumulated into a past due balance, so we asked AR for a repayment plan, and their answer was: wait. The credit outcome would affect what we owed, and they needed the case to resolve before we could finalize any repayment agreement.\n\nIn response to an email titled \"Seeking repayment plan for past due balance,\" an AR representative told us:\n\n> _Upon checking, the support case related to billing adjustment review via case ID <redacted> remains active. Please continue to monitor the support case for any update and provide the required information for faster resolution. Once the billing adjustment review for re-appeal was completed, we can discuss for possible payment plan._\n\nIn the meantime, we paid what we could and waited for resolution. Our balance stayed on the books for ten months\u2014and with it came past-due notices with threats of suspension we were told were automated and couldn't be stopped. Those are what prompted us to ask our account manager, more than once, whether we were truly at risk of suspension, and his answers gave us confidence that, difficult as the situation was, we were managing it the right way.\n\nWhen those warnings arrived, we did the same thing: made a payment and notified AR. That pattern\u2014always paying, always communicating\u2014put us squarely in the playbook our account manager described, and it seemed to work because we never appeared on his at-risk list.\n\n### A Balance We Couldn't Pin Down\n\nWhen the billing case finally resolved in January 2026, we expected to finalize a repayment plan, but we were blocked _again_\u2014and explaining why requires a brief note on how the approved credits worked:\n\nThe credits didn't arrive as a direct reduction to our outstanding balance. They came in eight different amounts, apparently as cash refunds\u2014separate transactions that made their way back to us through different payment methods. This meant our balance didn't shrink by the amount of the credits as we expected it would; and, more importantly, some of the credits remained outstanding due to two transactions we couldn't locate. Whether and how _those_ had been applied\u2014_unclear_.\n\nThose two missing refunds had no payment method attached in our billing console, and we couldn't find matching deposits in any of our bank or card accounts. They were material to calculating our actual outstanding balance, and any repayment plan, and without AR's help, we couldn't track them down.\n\nSo we submitted our first request for help on February 12:\n\n> _We are having trouble balancing our remaining amount due \\($\\[redacted\\]\\) against our credits awarded Jan8 \\($\\[redacted\\]\\). Can you help us track down these two \\(2\\) refund transactions? \\[...\\] We can't find a payment method associated with those two refunds \\[...\\] And we can't find matching transactions anywhere in any of our card or bank account statements. Do you know how / where those refunds were applied?_\n\nThree days later, on Feb 15, we received a response that simply restated the information in our billing console. It didn't answer our question.\n\nWithin an hour of receiving it, we responded and made the stakes explicit:\n\n> _Sorry if I wasn't clear: We have not been able to find these two refunds: \\[...\\] We cannot find the money. We believe we did not actually receive it. We have a general repayment plan in mind that we hope to present, but first we need to resolve this question of these two missing refunds. Without being able to account for that $\\[redacted\\], we believe our current outstanding balance could be off by that much. Your input on this is crucial for us._\n\nWe received no response.\n\nGrowing concerned about a possible disconnect with AR, that Friday, February 20, we cc'd our account manager on another follow-up to AR and, separately, forwarded the full thread to him directly, asking him to intervene:\n\n> _Is there any way you can help us get a response from \\[Redacted\\] or someone else in Billing regarding the missing refund payments? We want to implement a repayment plan, but need to be clear on our numbers._\n\nHe replied the same day:\n\n> _Yes \u2014 I'll reach out to her. \\[...\\] Let me pop into the ticket to see if I can escalate._\n\nThat was the last we heard from anyone.\n\nOn February 19, we received a past-due notice with a suspension warning dated the following day. This wasn't unusual\u2014we'd received them throughout the entire ten-month period and had always responded the same way: we made a payment and notified AR. We made a $10,000 payment on February 20 and let AR know. We interpreted the passing of February 20 the way we always had: as confirmation that our payment had been received and the threat had passed.\n\nFour days later, we were suspended.\n\n### Six Weeks Before Suspension\n\nSix weeks before our suspension, on January 14, we filed a second billing dispute. We had been trying to _reduce_ our hosting costs by shrinking our production database; instead, a mandatory storage configuration upgrade blocked us for 27 days and roughly _doubled_ our costs during that period. We disputed the incremental charges.\n\nAfter working through all required questions, our request was formally submitted to Billing's specialist team:\n\n> _Since these charges are not yet present on your account we here at Billing team cannot promise a refund or credits to offset these charges. However, we can connect with your AWS Account Manager for getting an exception approved for you by our specialist internal team. Please contact your Account Manager so they can work with the team for the required approval._\n\nWe formally looped in our AM as directed\u2014he was already cc'd on the case, but we followed the process\u2014and he confirmed he was actively monitoring and advocating for our credit request.\n\nOn January 22, we experienced a temporary account restriction\u2014we were unable to add certain services needed for our workaround. We reached out to our AM and sent a parallel update to our AR contact, informing her of the open billing case and our AM's involvement. The restriction was lifted.\n\nOur AR contact replied with direction nearly identical to what we'd received throughout our first billing dispute:\n\n> _Kindly continue to monitor the support case for billing adjustment review._\n\nThe technical workaround we selected took 20 days\u2014January 22 through February 10\u2014spanning our two most recent billing periods.\n\nWhen the effort was complete, we updated Billing, and our account manager, with the final timeline, cost, and the amount of our credit request. By February 12, everyone had the same information: the work was done, the dispute was submitted, and our last two invoices were under active dispute.\n\nBased on prior experience and similar direction from AR, we believed charges under active dispute would not be counted toward our outstanding balance\u2014let alone any balance used to justify suspension.\n\nWe don't know whether our disputed billing charges _were_ part of the outstanding balance that triggered suspension, or not. All we know is that AR knew about this open case, and so did our AM, who had told us more than once that he reviewed at-risk accounts before any suspension action was taken.\n\nNeither warned us that suspension was imminent.\n\n\u2014\n\nTo recap: we were paying\u2014over $100K during the same six months the internal note claimed we'd paid nothing. We were communicating\u2014multiple messages to AR in the twelve days before suspension, plus a direct escalation to our account manager after AR went silent. We had an open billing dispute, and a repayment plan in progress with a specific blocker we had been asking AR to help us resolve for nearly two weeks.\n\nNone of the criteria our account manager described for suspension applied to us. We hadn't stopped paying. We hadn't stopped communicating. We had never appeared on his at-risk list. The suspension hadn't had his sign-off or his manager's.\n\nOn the day of our suspension, our account manager reviewed our account and told us it \"shouldn't have happened.\"\n\nThree days later, when someone finally acted, service was restored within minutes.\n\nTo this day, we still don't know exactly why the suspension happened, or why it took three days to reverse.\n\n## Open Questions \\(Our Formal Request for Answers\\)\n\nOn March 4, three business days after our service was restored, we sent a formal request for answers to our hosting provider. As of this writing\u2014fourteen days later\u2014we have not received a response.\n\nWe're not publishing those questions verbatim here because we don't want to frame them as a public ultimatum. But our questions cover four areas:\n\nWe want to understand what triggered the suspension and who authorized it\u2014whether it was automated or a human decision, and why we were suspended when we had never appeared on our account manager's at-risk list.\n\nWe want to know why our billing console showed no payment records for a period in which we had paid over $104K in verified bank transactions, and what prompted the internal note.\n\nWe want to know why our missing refund issue\u2014raised six times between February 12 and February 24\u2014was never substantively addressed before we were suspended.\n\nAnd we want to understand why reinstatement took three days after we provided evidence of our payments, our open billing cases, and our multiple attempts to resolve the blocker to our repayment agreement\u2014especially when service was restored in _minutes_ once someone finally acted.\n\nWe're committed to updating this post-mortem when we have those answers.\n\n## What We've Learned\n\nTwo things we should have done differently, and one assumption we should never have made.\n\nThe process guidance our account manager provided was verbal. We relied on it in good faith\u2014it was specific, consistent, and came from someone with direct knowledge of our account. But we never sought written confirmation. We should have.\n\nWe noticed discrepancies in our billing console earlier in this process. We flagged them but didn't fully understand their significance, or their potential connection to how our account might be evaluated\u2014until it was too late to escalate effectively. That's a failure we must own.\n\nAnd finally: we assumed we were safe. We were doing everything we'd been told to do. The most recent suspension warning had passed without incident\u2014as others had. We interpreted that silence as confirmation. It wasn't. We should have insisted on written confirmation of our account status; and when we couldn't get AR to respond, we should have escalated beyond email and beyond our account manager.\n\n\u2014\n\nThat's our account of what happened. What follows is what we're doing about it.\n\n## What We're Doing\n\nThe most significant thing we've done to prevent a recurrence is reduce our monthly database hosting costs by approximately 50% from their peak. We accomplished this by completing a major database project we'd been working toward for over a year\u2014the same project that was blocked twice, producing the unplanned cost spikes at the center of both billing disputes. That pressure is gone. We can easily meet our regular charges going forward.\n\nWe have nearly paid off our outstanding balance\u2014by 50% if disputed charges are included, by more if not\u2014and we've already begun implementing our repayment plan. We met with our provider this week to formalize that, but we're not waiting to execute. If we prevail in our open billing dispute, the remaining balance shrinks further, making repayment that much easier and faster.\n\nWe are standing up failover infrastructure on a second provider, expected to be operational within 60 days. No single vendor should ever again have the ability to take our service offline with no recourse.\n\n### Going Forward\n\nGoing forward, we will not rely on verbal guidance or implicit signals to determine our account standing. We will seek explicit, written confirmation of our status from both our account manager and a dedicated AR contact\u2014someone we hope will commit, in writing, to direct communication \\(not system warnings\\) that gives us at least one opportunity to resolve any issue before adverse action is taken. We will also get in writing the exact steps our provider follows before suspending an account.\n\nFinally, we owe you a commitment about how we communicate during incidents. During this outage, there were moments\u2014particularly on Day 1\u2014where we knew more than we shared publicly. We held back because we didn't fully believe what was happening, and because we kept expecting it to be reversed at any moment. That was the wrong call. Going forward, we will share what we know as soon as we know it, even when it's embarrassing, scary, and we don't yet have the full picture. You deserve that transparency in real time, not just in a post-mortem.\n\n## Making it Right\n\nEvery paying customer affected by this outage will receive a credit equal to four days of their current subscription\u2014one day for each day of interrupted service. This credit will be applied automatically to your next bill.\n\nWe understand that's not enough\u2014it's an attempt to be equitable while not harming ourselves more than we've already been harmed by this outage. We want you to know that if anyone feels it doesn't properly compensate for how the outage affected you, just reach out and we will refund the whole month. No questions asked.\n\nTo our open-source users: the best we can offer is this post-mortem. We know some of you have been through similar things. If it's helpful for you to discuss further, please reach out. We'd also be grateful to hear from anyone who's been through something similar and what you did about it from a hosting perspective. And thanks for the HugOps. We're ready to share some back now.\n\nIf you have questions about anything in this post-mortem, we're at [ops@coveralls.io](mailto:ops@coveralls.io).\n\n## Acknowledgments\n\nFinally: thank you.\n\nDuring the worst week in our 15-year history, our community showed up. Customers offered server rack space. An SRE director offered to apply pressure through his own provider relationships. People we'd never met sent messages of support and solidarity.\n\nIn the DevOps community, the term is: HugOps. We received a lot of it that week, and it made a real difference.\n\nFor those who've stayed with us through this\u2014thank you. We know you have other options. We appreciate the opportunity to make this right.\n\n\u2014\n\n## Notes\n\n\\[1\\] _That ratio matters for this case. Our revenue doesn't scale proportionately to our infrastructure costs. Unexpected cost spikes hit us harder than they would a company whose hosting bill is covered by proportional revenue, which is part of why the billing situation at the center of this story developed the way it did._\n\n\\[2\\] _AWS RDS Extended Support Fees \\(ESF\\) are charged per vCPU per hour on any RDS instance running a PostgreSQL major version past its end of standard support date \u2014 on top of normal instance costs. Our production instance was a db.r6g.16xlarge: 64 vCPUs, 512 GB RAM, at the 65TB maximum RDS storage allocation. At 64 vCPUs and the Year 1 ESF rate, the surcharge approached the cost of the instance itself. Our Major Version Upgrade \u2014 which took four and a half months, three and a half of which fell within the Extended Support window \u2014 required a Blue/Green Deployment, meaning we ran two db.r6g.16xlarge instances simultaneously for that entire period, each accumulating ESF charges around the clock. Two instances at normal cost plus ESF on both: that is what nearly quadrupled our database hosting costs for the period._\n\n\\[3\\] _Note: In connection with the credit award, we signed an agreement that restricts us from disclosing specific amounts or characterizing our provider's position on the dispute._\n\n\u2014\n\n[Download a PDF version of this post-mortem](https://s3.amazonaws.com/assets.coveralls.io/statuspage/incidents/20260224/postmortem/Coveralls-Post-Mortem-Service-Outage-Feb-2026.pdf).",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-02-27T14:30:00.000-08:00",
"resolved_inferred": false,
"started_at": "2026-02-24T14:52:29.054-08:00",
"state": "postmortem",
"title": "Service outage (RESTORED, MONITORING)",
"updated_at": "2026-03-31T14:59:08.369-07:00",
"url": "https://stspg.io/1sfctfdlmr9c"
},
{
"body": "Just a note to address the gap in our Coverage Calculation Background Job Dequeue Time Graph today, MON, FEB 2 from 00:15:00 PST to 07:35:00 PST:\n\nCoveralls experienced no disruption in service at this time. Instead, a deployment issue cause the cron job that reports the metric to fail until it was resolved at 07:35:00 PST.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-02-02T09:33:52.000-08:00",
"resolved_inferred": false,
"started_at": "2026-02-02T09:33:51.977-08:00",
"state": "resolved",
"title": "All Systems Operational",
"updated_at": "2026-02-02T09:34:44.341-08:00",
"url": "https://stspg.io/gn1j89710t5h"
},
{
"body": "We had to pause some queues to recover performance for new jobs, but will clear those as soon as we've recovered normal build times for new jobs (ETA: 15-min).\n\nIf you are a user in APAC or EU, you may have had your jobs paused. One repo in particular has represented 90% overnight workload. We will reach out to that user.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-12-19T07:24:16.835-08:00",
"resolved_inferred": false,
"started_at": "2025-12-19T06:34:44.392-08:00",
"state": "resolved",
"title": "Elevated Latency in APAC and EU",
"updated_at": "2025-12-19T07:24:16.861-08:00",
"url": "https://stspg.io/02l24lf4x109"
},
{
"body": "Build times are normal for all new builds. Background queues have been cleared of 99% of backlogged jobs; the only ones that remain are for large repos (>5K source files), which should clear in the next 30-45 minutes depending on size.\n\nNo further effects on latency expected. \n\nClosing this incident.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-12-18T08:58:27.156-08:00",
"resolved_inferred": false,
"started_at": "2025-12-18T06:51:18.919-08:00",
"state": "resolved",
"title": "Elevated Latency in APAC and EU",
"updated_at": "2025-12-18T08:58:27.172-08:00",
"url": "https://stspg.io/7hg0tqpk0948"
},
{
"body": "This incident has been resolved. Build times for all new builds is normal across the board.\n\nWe are still clearing a backlog of background jobs, which should be clear in the next 15-20-min.\n\nIf you are having issues with a slow or stuck build, feel free to reach out to us at support@coveralls.io. These steps will save time:\n\n1) Mention this incident:\nhttps://status.coveralls.io/incidents/hkqt790213m5\n\n2) Share your Coveralls Build URL (from your CI build log), or your CI build number.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-11-17T11:43:38.674-08:00",
"resolved_inferred": false,
"started_at": "2025-11-17T08:35:03.565-08:00",
"state": "resolved",
"title": "Elevated Latency",
"updated_at": "2025-11-17T11:43:38.694-08:00",
"url": "https://stspg.io/79r8vvnz4l3m"
},
{
"body": "We have attempted to improve latency this morning for new jobs from users in EU and US time zones, which has meant offloading older jobs to specialized queues with scaled up resources, which, at this point, have already drained.\n\nSystem-wide, latency continues improving for all users and should reach normal in 15-30 minutes for most users.\n\nIf you are still experiencing elevated latency after that, or have any jobs from CI builds run in the last 24 hrs that have not completed, please reach out to us at support@coveralls.io and we'll investigate to determine whether your builds were caught up in this incident, or if they have a different cause.\n\nNOTE: Missing data points in the \"DEQUEUE\" Graph on our main Status Page does not indicate that processing stopped during that period, just that the server that reports our stats lost comms during that period. As stated elsewhere, to avoid this confusion in the future it's one of our goals to spread stats collection across all servers.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-11-10T10:55:54.000-08:00",
"resolved_inferred": false,
"started_at": "2025-11-10T08:14:00.004-08:00",
"state": "resolved",
"title": "Elevated Latency",
"updated_at": "2025-11-10T11:03:51.760-08:00",
"url": "https://stspg.io/j214kcyh341b"
},
{
"body": "We are closing this incident having received no further reports of 504 errors today. We will continue to monitor for them.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-10-24T14:54:04.600-07:00",
"resolved_inferred": false,
"started_at": "2025-10-24T11:35:11.000-07:00",
"state": "resolved",
"title": "504 Gateway Timeouts (Resolved)",
"updated_at": "2025-10-24T14:54:04.614-07:00",
"url": "https://stspg.io/b71rcnbr99wp"
},
{
"body": "We have received no further reports of 504 timeout errors on coverage report uploads today, but we continue to monitor and will continue trying to improve our mitigations.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-10-13T16:15:42.403-07:00",
"resolved_inferred": false,
"started_at": "2025-10-13T11:02:40.474-07:00",
"state": "resolved",
"title": "Some reports of 504 Timeouts",
"updated_at": "2025-10-13T16:15:42.420-07:00",
"url": "https://stspg.io/3424pqnymqnp"
},
{
"body": "This incident is resolved. \n\nWe have applied an additional layer of monitoring that should help us catch these cases earlier.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-10-10T12:00:40.425-07:00",
"resolved_inferred": false,
"started_at": "2025-10-10T08:23:27.445-07:00",
"state": "resolved",
"title": "504 Timeouts (Resolved)",
"updated_at": "2025-10-10T12:00:40.440-07:00",
"url": "https://stspg.io/fjs06cmd8l4d"
},
{
"body": "We are closing this incident after recent mitigations and a weekend without any reports.\n\nWe continue to implement mitigations and infrastructure changes we believe will further reduce incidents of this error type.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-09-29T08:56:14.000-07:00",
"resolved_inferred": false,
"started_at": "2025-09-26T08:28:03.000-07:00",
"state": "resolved",
"title": "504 Timeout Errors on Coverage Uploads",
"updated_at": "2025-09-29T08:56:55.974-07:00",
"url": "https://stspg.io/km3p38qs5n00"
},
{
"body": "**This is a postmortem on this specific issue:** Intermittent 500 Errors on Coverage Uploads\n\n**Summary**  \nBetween September 20\u201324, some customers experienced intermittent `500 Internal Server Error` responses during coverage uploads \\(`POST /api/v1/jobs`\\). The issue was initially hard to diagnose because:\n\n* Failures did not surface reliably in our error tracker \\(BugSnag\\).\n* They appeared to affect only some requests, some customers.\n\n**Impact**\n\n* Some coverage uploads failed to process, causing build reporting delays or gaps.\n* Frequency was low enough to appear intermittent, which delayed detection and resolution.\n\n**Timeline**\n\n* **Sep 20\u201323**: First customer reports of intermittent 500s. Initial theories involved a regression in a recent release of our coverage-reporter integration \\(client-side\\).\n* **Sep 23\u201324**: Deep log analysis across ELB and application logs revealed errors concentrated on a single web server.\n* **Sep 24**: Confirmed that server alone was responsible for thousands of SSL-related failures \\(`Faraday::SSLError`, `OpenSSL::SSL::SSLError`, `Seahorse::Client::NetworkingError`\\). Other servers were clean.\n* **Sep 24**: Mitigation: that server was destroyed. Errors ceased immediately.\n\n**Root Cause**  \nIn terms of possible cause, we believe this additional Web server was provisioned during autoscaling with a different Ubuntu version than the rest of the fleet. This seemed to result in a broken or outdated CA certificate store, causing outbound SSL connections \\(GitHub, Travis, etc.\\) to fail intermittently\u2014but then bubble up to a `500` error for the original request \\(`POST` to `/api.v1/jobs`\\).\n\n**Resolution**\n\n* Problematic server removed from service.\n* Future mitigation: verify baseline OS/version and CA store when adding new servers, especially via automation.\n* Next step: document the correct procedure to disable a single server in Cloud66 load balancers, instead of outright destroying, so we can retain the server for forensic investigations.\n\n**Lessons Learned**\n\n* Errors can hide if they don\u2019t surface in the bug tracker. Direct log analysis is essential.\n* Even one misconfigured server can cause significant customer impact.\n* Consistency in OS/base image, and CA, is critical.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-09-24T17:53:00.000-07:00",
"resolved_inferred": false,
"started_at": "2025-09-03T11:00:31.000-07:00",
"state": "postmortem",
"title": "Elevated 504 Timeout Errors",
"updated_at": "2025-09-25T09:41:05.993-07:00",
"url": "https://stspg.io/lhyfrgjdj1dj"
},
{
"body": "We believe this issue has been resolved for now.\n\nThe underlying cause still appears to be large spikes in incoming Web traffic from other outlier repositories that we have not yet identified or not yet paused.\n\n**Interim Solution**:\n\nTo reduce the risk of recurrence, we have applied _temporary load balancer adjustments_ that change how requests are distributed, which should _lower_\u2014if not _eliminate_\u2014the frequency of **503** \u201c**This website is under heavy load**\u201d **errors**.\n\n**Permanent Solution**:\n\nWe are also designing a permanent solution to _rate-limit abnormal request patterns_. This will require coordination at the policy/SLA level before it can be fully implemented.\n\nIn the meantime, we will continue to closely monitor traffic and use targeted load balancer and web server configurations to mitigate the impact of outlier traffic spikes.\n\n**More details**:\n\nFor a more detailed assessment / RCA of this incident and its recent, related incidents, see [this postmortem](https://status.coveralls.io/incidents/wqbsxnzv0jsf).\n\n**Update \\(Thu, Aug 21\\)**:\n\nWe have identified a **different permanent solution**, which does not entail changes to SLA-level details for \u201coutlier repos.\u201d We may still implement such a solution, but our alternative solution should avoid further 503 errors and be implemented in the next 48-72 hrs.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-08-20T15:21:47.849-07:00",
"resolved_inferred": false,
"started_at": "2025-08-20T08:48:22.100-07:00",
"state": "postmortem",
"title": "More reports of \u201cWebsite under heavy load\u201d",
"updated_at": "2025-08-21T10:29:55.996-07:00",
"url": "https://stspg.io/4jcpgc1b166l"
},
{
"body": "### Postmortem: Reports of \u201cWebsite under heavy load\u201d errors\n\nWe experienced multiple intermittent errors over the past several days before we were able to identify the true root cause and resolve the issue.\n\n**Root Cause**  \nThe errors were caused by a single outlier repository generating extremely high-volume requests \\(750\u20131,800\\+ coverage report uploads per build\\). Combined with the default \u201csticky request\u201d behavior in Passenger Enterprise \\(which routes repeat requests from the same IP to the same HTTP server\\), this overwhelmed individual servers. Once a server\u2019s request queue was exhausted, subsequent requests returned a `503` error with the message: _\u201cThis website is under heavy load.\u201d_\n\nAlthough each server was able to process individual requests within normal timeframes, the concentrated traffic volume from a single repo and source IP could not be evenly distributed across servers. This led to repeated saturation of request queues and customer-visible errors.\n\n**Solutions Implemented**\n\n1. We are testing new settings to override Passenger\u2019s default \u201csticky request\u201d behavior to allow requests to be distributed more evenly across servers.\n2. We have paused processing for the outlier repository while we validate that the configuration changes are sufficient to prevent future incidents.\n\n**Next Steps**\n\n* Continue monitoring system performance to confirm stability.\n* Reintroduce the paused repository once we are confident the mitigations are effective.\n\n**Closing**  \nWe appreciate your patience as we worked through this issue. These changes are intended to permanently guard against similar incidents going forward. If you encounter unexpected errors, please contact us at [support@coveralls.io](mailto:support@coveralls.io).\n\n**Related incidents**\n\n1. **Aug 13**: [Intermittent request rejections](https://status.coveralls.io/incidents/1n7plxrj8j44)\n2. **Aug 14**: [Service unavailable with HTML error page or 500 errors](https://status.coveralls.io/incidents/v5mcbrsbhgt4)\n3. **Aug 18**: [Reports of \"Website under heavy load\" errors](https://status.coveralls.io/incidents/fr6sp5kyn128)\n4. **Aug 19 \\(Today\\)**: [Reports of \"Website under heavy load\" errors](https://status.coveralls.io/incidents/wqbsxnzv0jsf)",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-08-19T10:26:33.373-07:00",
"resolved_inferred": false,
"started_at": "2025-08-19T08:21:44.595-07:00",
"state": "postmortem",
"title": "Reports of \"Website under heavy load\" errors",
"updated_at": "2025-08-19T12:00:50.565-07:00",
"url": "https://stspg.io/3srf5qdsc4s1"
},
{
"body": "A postmortem for this incident and its related incidents has been posted [here](https://status.coveralls.io/incidents/wqbsxnzv0jsf):\n\n* [**Postmortem: Reports of \u201cWebsite under heavy load\u201d errors**](https://status.coveralls.io/incidents/wqbsxnzv0jsf)\n\n**Related incidents**\n\n1. **Aug 13**: [Intermittent request rejections](https://status.coveralls.io/incidents/1n7plxrj8j44)\n2. **Aug 14**: [Service unavailable with HTML error page or 500 errors](https://status.coveralls.io/incidents/v5mcbrsbhgt4)\n3. **Aug 18**: [Reports of \"Website under heavy load\" errors](https://status.coveralls.io/incidents/fr6sp5kyn128)\n4. **Aug 19 \\(Today\\)**: [Reports of \"Website under heavy load\" errors](https://status.coveralls.io/incidents/wqbsxnzv0jsf)",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-08-19T01:10:17.588-07:00",
"resolved_inferred": false,
"started_at": "2025-08-18T12:25:56.000-07:00",
"state": "postmortem",
"title": "Reports of \"Website under heavy load\" errors",
"updated_at": "2025-08-19T12:00:55.412-07:00",
"url": "https://stspg.io/xyv9nm9f08k9"
},
{
"body": "A postmortem for this incident and its related incidents has been posted [here](https://status.coveralls.io/incidents/wqbsxnzv0jsf):\n\n* [**Postmortem: Reports of \u201cWebsite under heavy load\u201d errors**](https://status.coveralls.io/incidents/wqbsxnzv0jsf)\n\n**Related incidents**\n\n1. **Aug 13**: [Intermittent request rejections](https://status.coveralls.io/incidents/1n7plxrj8j44)\n2. **Aug 14**: [Service unavailable with HTML error page or 500 errors](https://status.coveralls.io/incidents/v5mcbrsbhgt4)\n3. **Aug 18**: [Reports of \"Website under heavy load\" errors](https://status.coveralls.io/incidents/fr6sp5kyn128)\n4. **Aug 19 \\(Today\\)**: [Reports of \"Website under heavy load\" errors](https://status.coveralls.io/incidents/wqbsxnzv0jsf)\n\n\u200c\n\n**Previous postmortem, prior to final resolution:**\n\nWe received customer reports early this morning PDT, Thu, Aug 14, indicating that the [Coveralls.io](http://Coveralls.io) service was unavailable for a substantial period overnight.\n\nWhile our Dequeue Graph appears unresponsive since 18:00 UTC yesterday Aug 13 / 11:00 PDT yesterday, Aug 13, this does not seem to be related to the unavailable period, which appears to have started around 02:00 UTC Aug 14 / 22:00 PDT Aug 13, because we were in service all day and afternoon yesterday, Wed, Aug 13, through EOD PDT.\n\nOur uptime ping did not catch this because Phusion Passenger was returning an HTML error page. Some customers report receiving a 500 error in that message; others report simply receiving HTML as their API response instead of JSON.\n\nWhile we have not seen production bugs specific to the issue, we suspect it\u2019s because the issue occurred at the Passenger/Nginx layer, which simply prevented new traffic since occurring. As far as previously received coverage report uploads, we see successful background processing of jobs in queue all through the evening into this morning.\n\nWe suspect as root cause a regression from a deploy around 5pm yesterday, due mainly to the fact that rolling back to a previous deploy has, so far, resolved the issue.\n\nWe will investigate further to obtain a full understanding of the root cause, but in the meantime, we\u2019ll continue monitoring.\n\nPlease reach out if you are still experiencing any form of error page in Web requests, or any non-200 API response from C[overalls.io.](http://Coveralls.io)\n\nOur apologies for the inconvenience.\n\n**Note**: **Recommended** **Best Practice**: `continue-on-error: true`\n\nAs a general best practice, we recommend adding the **input option** `continue-on-error: true` \\(for [Coveralls GitHub Action](https://github.com/marketplace/actions/coveralls-github-action#inputs)\\), or `continue_on_error: true` \\(for [Coveralls Orb for CircleCI](https://circleci.com/developer/orbs/orb/coveralls/coveralls#commands-upload)\\); or the `--no-fail` **flag** \\(for [Coveralls Universal Coverage Reporter](https://github.com/coverallsapp/coverage-reporter?tab=readme-ov-file#usage)\\), which will keep Coveralls errors from breaking your CI workflows in the future.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-08-14T07:45:56.622-07:00",
"resolved_inferred": false,
"started_at": "2025-08-14T07:35:14.384-07:00",
"state": "postmortem",
"title": "Service unavailable with HTML error page or 500 errors",
"updated_at": "2025-08-19T11:59:29.399-07:00",
"url": "https://stspg.io/h69wjdsj5hsr"
},
{
"body": "A postmortem for this incident and its related incidents has been posted [here](https://status.coveralls.io/incidents/wqbsxnzv0jsf):\n\n* [**Postmortem: Reports of \u201cWebsite under heavy load\u201d errors**](https://status.coveralls.io/incidents/wqbsxnzv0jsf)\n\n**Related incidents**\n\n1. **Aug 13**: [Intermittent request rejections](https://status.coveralls.io/incidents/1n7plxrj8j44)\n2. **Aug 14**: [Service unavailable with HTML error page or 500 errors](https://status.coveralls.io/incidents/v5mcbrsbhgt4)\n3. **Aug 18**: [Reports of \"Website under heavy load\" errors](https://status.coveralls.io/incidents/fr6sp5kyn128)\n4. **Aug 19 \\(Today\\)**: [Reports of \"Website under heavy load\" errors](https://status.coveralls.io/incidents/wqbsxnzv0jsf)",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-08-13T12:24:03.000-07:00",
"resolved_inferred": false,
"started_at": "2025-08-13T10:14:07.058-07:00",
"state": "postmortem",
"title": "Intermittent request rejections",
"updated_at": "2025-08-19T12:02:53.441-07:00",
"url": "https://stspg.io/qbz08fdtjbjt"
}
]
}