{
"vendor": "TeraSwitch",
"slug": "teraswitch",
"platform": "statuspage",
"status_url": "https://www.teraswitchstatus.com",
"last_checked": "2026-09-16T12:28:20Z",
"last_state": "ok",
"history_backfilled": true,
"first_watched": "2026-09-04T07:06:16Z",
"incidents": [
{
"body": "All links have been restored and are stable. Traffic is fully normalized.",
"first_seen": "2026-09-04T14:19:03Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-09-06T05:52:37.249Z",
"resolved_inferred": false,
"started_at": "2026-09-04T13:09:58.363Z",
"state": "resolved",
"title": "Traffic detour - LON/DUB/EWR/NY undersea fiber cable offline",
"updated_at": "2026-09-06T05:52:37.264Z",
"url": "https://stspg.io/blc9h3xcyb5k"
},
{
"body": "AWS has restored normal operation in the TYO market.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-09-03T20:04:22.999Z",
"resolved_inferred": false,
"started_at": "2026-09-03T14:52:14.705Z",
"state": "resolved",
"title": "AWS Tokyo - Ongoing AWS traffic detours",
"updated_at": "2026-09-03T20:04:23.012Z",
"url": "https://stspg.io/7q02723fqk1z"
},
{
"body": "The connection between NY and Chicago has been repaired and has remained stable for multiple hours.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-08-22T14:27:19.240Z",
"resolved_inferred": false,
"started_at": "2026-08-22T02:45:31.045Z",
"state": "resolved",
"title": "EWR/CHI - Fiber cut and Network Reroute",
"updated_at": "2026-08-22T14:27:19.254Z",
"url": "https://stspg.io/2913538rjj9s"
},
{
"body": "_\\(A more detailed/illustrated RCA in PDF format is available on request\\)_\n\n# Internet Outage Event at Multiple Sites\n\n**Incident date:** August 12, 2026 \u00a0\u00b7\u00a0 **Regions affected:** EU, APAC \u00a0\u00b7\u00a0 **Full recovery:** 33 minutes\n\n## Executive summary\n\nIn the early morning of August 12, 2026 \\(UTC\\), customers at thirteen of Teraswitch's 31 sites lost reachability to Internet destinations and to other Teraswitch sites over the private backbone: twelve across the EU and APAC regions, and MIA1 in North America, which was intentionally and temporarily removed from the backbone as a containment measure. The cause was a stale default route \\(`0.0.0.0/0`\\) originated by an MIA1 edge router: a static route left over from the site's turn-up, redistributed into BGP. The static is deliberately less preferred than any transit-learned default and had sat inactive for months, until a routine transit provider maintenance withdrew the local external default shortly before the first alarm and activated it. As originated, the route carried an empty AS path containing no external networks, a metric of 0, and no communities: the ordinary signature of a redistributed static route.\n\nA digit transposition in a route-map on MIA1's sessions to the global route reflectors, `20326:301:x` written where `20326:130:x` was intended, then placed the route into the community lists that control regional export, and the global reflectors added the no-export community to the copies sent toward the EU and APAC regions. The route reflector hierarchy propagated it to our EU and APAC markets carrying that flag.\n\nBecause the attributes that normally distinguish a remote default from a local one were never attached, receiving edge routers at the affected sites evaluated the route as locally originated and preferred it over their own valid local default. No-export is a control flag that forbids a route from being advertised across a BGP boundary into another ASN. The connection from the edge routers to the data center fabric is exactly such a boundary. When the affected edge routers selected the stale route as their best default, no-export prevented them from advertising any default to their spines. With no default route present, the affected fabrics stopped forwarding Internet-bound traffic, isolating them from their own healthy edge routers.\n\nThe first alarm fired at `03:49:52 UTC`. Engineers identified the stale route within 10 minutes; the remainder of the outage window was spent executing containment, removing MIA1 from the backbone and powering off the AMS2 route reflector, and waiting for routes to recalculate. Reconvergence began at `04:16:15 UTC`, with service returning progressively as sites withdrew the stale default and reconverged on their local default routes; routes fully recalculated and traffic returned to normal 33 minutes after the first alarm. A global configuration change deployed to all compute sites later the same day ensures that a stale or invalid route entering the fabric can no longer prevent traffic forwarding.\n\n**Sites affected:** 13 of 31 \u00b7 **Time to identify:** <10 min \u00b7 **Recovery begins \\(UTC\\):** 04:16:15 \u00b7 **Alarm to full recovery:** 33 min\n\n## Customer impact\n\nCustomers at **LON1, AMS1, AMS2, AMS3, DUB1, DUB2, FRA2, SGP1, SGP2, TYO1, TYO2 and TYO3** experienced loss of reachability to Internet destinations, as well as to internal Teraswitch backbone destinations between affected sites.\n\nAffected workloads included latency and availability sensitive blockchain infrastructure hosted in these regions, including Solana validator and RPC nodes, which lost Internet and inter-site reachability for the duration of the event. We understand how sensitive and how large these workloads are, and the level of trust their operators place in us.\n\n### Impacts under active remediation\n\nCustomer DoubleZero's GRE/VPN tunnel sessions were severed when Bare Metal services became unable to reach the DoubleZero devices. We are working to move these interconnects to within the data center fabric, removing their dependency on paths this event was able to break.\n\nCustomer load balancers and servers remained able to receive inbound traffic: this was an outbound blackhole, affecting traffic leaving the fabric toward the Internet. Customers who announce their subnets to us via BGP therefore continued to send those announcements, attracting requests into data centers that could not respond to them. This is addressed by the overall hardening, which enforces that the fabric can always reach the Internet whenever its edge routers are otherwise healthy and attracting inbound traffic.\n\n### Not affected\n\nNorth American data centers were not affected.\n\nMIA1 was deliberately taken off the backbone as a containment measure. MIA1 is a small site, and Teraswitch engineering chose to drop it from the backbone quickly to resolve the overall issue. The impact to MIA1 operation was not monitored during this time, but some impact to connectivity at the site would be expected as it moved into island mode, with Internet routing taking local transit only. The site remained in service and was reintegrated into the backbone via a maintenance shortly after the event.\n\n## Timeline of events\n\nAll times UTC, August 12, 2026.\n\n* **~03:30** A routine transit provider maintenance begins. At an unrecorded moment during this window, likely shortly before the first alarm, er100a.mia1's external default route is withdrawn; the leftover static default activates, enters BGP carrying the transposed community values, receives no-export toward EU and APAC at the global reflectors, and begins propagating.\n* **03:49:52** First alarm fires: a monitoring host in LAX1 reports it cannot reach a host in LON1. The stale default route from MIA1 has propagated into EU and APAC, and affected site fabrics stop forwarding Internet traffic.\n* **\\+10 min from onset** Stale MIA1 default identified as the cause. The responding engineer acted immediately, shutting down the loopback interface carrying MIA1's SR-MPLS/IS-IS backbone connectivity and powering off the AMS2 route reflector, before reporting the cause. The exact moments were not recorded, but fall within this window; the remaining minutes before restoration were consumed by these containment actions and by route recalculation across the affected sites. In the same window the team halted route exports from the global route reflectors. The only operational change from this was peering traffic taking local and regional paths rather than crossing regions; customer and internal routes were maintained by our Continuity Reflectors, which handle only internal and customer routes.\n* **04:16:15** Reconvergence begins. With the AMS2 reflector shut down, the backlog of route changes was large, and withdrawing and reprogramming the stale default across the affected fabrics took time. Service returns progressively as each site reconverges on its local default routes.\n* **~04:23** Routes fully recalculated and traffic patterns return to normal. Full recovery, 33 minutes after the first alarm.\n* **05:18** Global configuration fix decided in coordination with our network hardware vendor; deployment to all compute sites begins.\n* **by 15:00** Global fabric hardening deployment completed across all compute sites, including unaffected data centers \\(before 11:00 AM Eastern US\\). The change was applied one edge router at a time, across the edge routers feeding approximately 62 spines.\n* **18:29** Incident marked resolved. Investigation of the underlying condition continues.\n\n## Background: how default route signaling works\n\nEach Teraswitch compute site contains a minimum of two Internet/backbone edge routers and two spine routers; larger sites run more of each \\(four edge routers and six spines, for example\\). Every edge router connects to every spine, and the spines connect to every leaf switch pair in the data center. Customer bare metal servers attach to the leaf layer.\n\nUnder the design in place at the time of the event, the edge routers shared the default route \\(`0.0.0.0/0`\\) they had themselves selected, from local transit or the backbone, into `vrf INTERNET` on the spines. This default route is the signal by which the leaf/spine network learns that Internet forwarding is available through its local edge routers. It is also the only route the edges send toward the fabric: the spines carry no more specific Internet or inter-site routes, so all traffic leaving a site, whether bound for the Internet or for another Teraswitch data center, follows this single default, and losing it severs both at once. When the local edges originate this export, the next hop is the edge routers themselves, addresses the spines can always resolve directly. Route attributes, specifically communities and metric, distinguish a default originated by the local site edges from one learned from a remote site over the backbone. Each site normally prefers the default originated by its own local edge routers.\n\nSites learn default routes from other backbone sites by design. If a site loses its own local default, meaning its local edge routers can no longer reach the Internet, the intended behavior is to carry traffic across the backbone to another Teraswitch site as quickly as possible, where healthy edge routers are likely available. The remote defaults exist as that escape path, and the distinguishing attributes exist so the escape path is only chosen when the local path is gone.\n\nRoutes move between sites through a redundant route reflection hierarchy: every site normally runs two local route reflectors \\(MIA1 temporarily did not, as described in the root cause analysis\\), every region \\(EU, APAC, US East, US West\\) has at least two regional reflectors, and three global reflectors tie the regions together. AMS2 hosts one of the EU regional reflectors. Alongside this hierarchy run the Continuity Reflectors: simple route reflectors that hold only Teraswitch data center prefixes and customer prefixes, at a lower priority than the normal reflector fabric, so that a total loss of the normal route reflectors does not cause major route instability or traffic blackholes. As described later, the hierarchy's redundancy could not mitigate this event, because the failure was one of route selection rather than reflector availability: a stale default that appears locally originated wins best path regardless of how many healthy alternatives exist.\n\n### Why the default was shared rather than unconditionally originated\n\nThis design was deliberate. By sharing a default route the edge had actually selected, rather than always originating one, an edge router could not come online and attract traffic from the spines before it held Internet routes and a usable default of its own. Without this logic, an edge router returning from maintenance or boot could blackhole traffic. The approach served traffic sequencing during maintenance and was believed to provide higher safety and availability.\n\n## What happened\n\nA default route originated by MIA1 edge router er100a.mia1 was propagated with an empty AS path containing no external networks \\(downstream devices in other ASNs would see only `AS20326`\\), a metric of 0, no communities of its own, and a next hop pointing at MIA1 router addressing. None of this was stripped in transit: these are the ordinary attributes of a static route redistributed into BGP. The source was a static default route configured during MIA1's turn-up, when the site had a single transit provider; it should have been removed when the second transit was activated, and was not. By design, our edge routers prefer transit-learned defaults over static ones, so the static sat inactive in the background for months, present in configuration but not in use. On August 12, a routine transit provider maintenance beginning around 03:30 UTC withdrew er100a.mia1's external default; the exact moment was not recorded, but the route change likely occurred shortly before the first alarm at 03:49:52 UTC. The static activated and entered BGP. The BGP policy between the two MIA1 edge routers intentionally does not share static routes, so er100b.mia1 never accepted the static default; it continued feeding both MIA1 spines its own working external default, which is why MIA1 itself initially stayed online.\n\nOn its way into the route reflector hierarchy, the route was stamped by an inbound policy on MIA1's direct sessions to the global reflectors. That policy contained a digit transposition, `20326:301:x` written where `20326:130:x` was intended, and the transposed values landed in the community lists that instruct the global reflectors to add the well-known `no-export` community toward specific regions. The copies exported toward EU and APAC therefore carried no-export; the copy distributed within North America did not, because the corresponding NA value \\(`20326:301:10`\\) was never written by the transposition. This is why the impact was confined to the EU and APAC markets. Within North America the route was present without the flag. Edge routers are largely indifferent to which default wins, and with no flag blocking the edge to spine advertisement of the MIA1 default toward the spine layer, the NA fabrics kept operating as normal. No-export is inert inside our own `AS20326`, so the route traveled the reflector hierarchy untouched: EU sites received it from the AMS2 regional reflector, APAC sites through the global reflector tier and the APAC regional reflectors. At each affected edge router it won best path, and because the edge to spine connection crosses a BGP boundary into a different ASN, no-export forbade advertising any default across it. The affected fabrics lost their only default.\n\n### Failure sequence at each affected site\n\n1. The site's edge routers receive the stale `0.0.0.0/0` from their upstream route reflectors.\n2. Because the standard communities that normally identify a remote default were absent, nothing demoted the route on arrival. With local preference equal, best path selection reached the AS path length comparison, where the route's empty AS path \\(length zero, shown in captures as only the origin code `?`\\) beat the local transit-learned default carrying its provider's ASN. The edge routers selected it over their own valid local default.\n3. The route carries the no-export community. The session from the edge routers toward `vrf INTERNET` on the spine layer crosses a BGP boundary into a different ASN, and no-export forbids the route from being advertised across it.\n4. The edge routers therefore stop advertising any default route to the spines. The fabric's only Internet default is withdrawn.\n5. With no acceptable default present in `vrf INTERNET`, the leaf/spine fabric stops forwarding Internet-bound traffic. The fabric is isolated from its own healthy edge routers.\n\n## Diagnosis and response\n\nThe initial picture was ambiguous. Not all sites were affected, and the pattern appeared geographic, which first pointed the responding engineers toward a DDoS event or a failure of DDoS mitigation systems. Time was also spent investigating whether IS-IS, our internal routing protocol, was involved: roughly a week earlier, a five minute instability had affected the backbone, and a recurrence was a plausible explanation. That earlier event proved to be unrelated. All of this took place within roughly the first ten minutes of the response.\n\n### Misleading Spine <-> Edge Traffic\n\nTwo observations made the true failure harder to see. First, traffic was still visible flowing from the edge layer to the spine layer at affected sites. The key detail, recognized later, was that this was L2 transport and private VRF traffic. Only the `INTERNET` vrf had failed, and the presence of healthy traffic on the same links delayed a closer look at the fabric's Internet routing specifically.\n\nSecond, management access looked completely normal, for reasons described below.\n\n### Management network behavior during the outage\n\nTeraswitch maintains a completely separate management network at each site, fronted by a high availability firewall pair. Under normal conditions these firewalls use the Teraswitch backbone, but they also connect to our serial console servers, which carry a dedicated and isolated Internet connection that is entirely separate from Teraswitch operation. When the data center networks became unusable, the firewalls failed over to this out of band connectivity automatically.\n\nThis design worked exactly as intended: our team had unfettered access to every device throughout the outage, as though everything were normal. For the same reason, it was slightly misleading. Because access felt normal, it appeared that nothing major was wrong with the equipment or the fabrics, when in fact the management plane had instantly routed around the very failure under investigation. Good for resolution and access, but misleading for diagnosis.\n\n### The moment of identification\n\nDespite the emergency response not recording the broken state of the network, the responding engineer captured the routing state on the AMS3 edge at the moment they located the issue:\n\n```\ner111a.ams3.teraswitch.com#show ip route 0.0.0.0/0\n\nVRF: default\nSource Codes:\n       C - connected, S - static, K - kernel,\n       O - OSPF, IA - OSPF inter area, E1 - OSPF external type 1,\n       E2 - OSPF external type 2, N1 - OSPF NSSA external type 1,\n       N2 - OSPF NSSA external type2, B - Other BGP Routes,\n       B I - iBGP, B E - eBGP, R - RIP, I L1 - IS-IS level 1,\n       I L2 - IS-IS level 2, O3 - OSPFv3, A B - BGP Aggregate,\n       A O - OSPF Summary, NG - Nexthop Group Static Route,\n       V - VXLAN Control Service, M - Martian,\n       DH - DHCP client installed default route,\n       DP - Dynamic Policy Route, L - VRF Leaked,\n       G  - gRIBI, RC - Route Cache Route,\n       CL - CBF Leaked Route\n\nGateway of last resort:\n B I      0.0.0.0/0 [200/0]\n           via 100.96.49.1/32, IS-IS SR tunnel index 103\n              via TI-LFA tunnel index 40, label 800961\n                 via 64.130.60.122, Ethernet25/1, label imp-null(3)\n                 backup via 64.130.60.81, Ethernet26/1, label 974060\n\ner111a.ams3.teraswitch.com#show bgp ipv4 unicast\nBGP routing table information for VRF default\nRouter identifier 100.124.63.2, local AS number 20326\nRoute status codes: s - suppressed contributor, * - valid, > - active, E - ECMP head, e - ECMP\n                    S - Stale, c - Contributing to ECMP, b - backup, L - labeled-unicast\n                    % - Pending best path selection\nOrigin codes: i - IGP, e - EGP, ? - incomplete\nRPKI Origin Validation codes: V - valid, I - invalid, U - unknown\nAS Path Attributes: Or-ID - Originator ID, C-LST - Cluster List, LL Nexthop - Link Local Nexthop\n\n          Network                Next Hop              Metric  AIGP       LocPref Weight  Path\n * >      0.0.0.0/0              100.96.49.1           0       -          100     0       ? Or-ID: 100.96.63.1 C-LST: 20.32.6.119 20.32.6.1\n```\n\nThis is the stale default as received at an affected edge: origin incomplete \\(the marker of a redistributed route\\), an AS path containing no external networks \\(making it appear locally originated\\), a next hop of `100.96.49.1` in MIA1 addressing, and reflection metadata, originator ID `100.96.63.1` with the AMS2 cluster list, showing the path the route took. The edge itself could resolve that next hop over an IS-IS SR tunnel across the backbone, so from the edge's perspective the route was usable. What made it unusable as the fabric signal was the no-export community it carried: once selected as best, it could not be advertised across the BGP boundary to the spines. Next-hop-self is applied on the edge to spine sessions, so next hop reachability played no part in the fabric failure.\n\n## Root cause analysis\n\n### Confirmed observations\n\n* The default route was received by EU and APAC edge routers without the standard communities that identify a legitimate remote default, and with an empty AS path containing no external networks: as originated, the bare signature of a locally redistributed static route. As received, it also carried the transposed large-community values described below and the no-export community, established by configuration analysis and lab replication of the full issue.\n* The route carried the originator ID of MIA1 edge router er100a.mia1 \\(`100.96.63.1`\\) and a cluster list of `20.32.6.119 20.32.6.1`: the AMS2 regional reflector and a global reflector, most recent reflector listed first. This matches MIA1's reflector topology at the time. MIA1's edge routers were connected directly to the global reflectors, a valid but intentionally temporary arrangement used because the edge routers were installed before the site's own reflector units, so no MIA1-local or North American regional entries would be expected on this path.\n* At the time of the event, er100b.mia1 was observed properly learning its external default route, and both MIA1 edge routers were receiving defaults from our other regional North American route reflectors. The stale static default sat inactive on er100a.mia1 by design, activating and entering BGP only when er100a.mia1's own external default was withdrawn during the provider maintenance that began around 03:30 UTC.\n* Receiving edge routers preferred the route because the attributes that distinguish local origination from remote origination were absent.\n* The route carried the no-export community, applied by the global reflectors' regional export policies through the transposed community values. Once an edge router selected this route as its best default, no-export forbade advertising any default across the BGP boundary to the spines, which under the then-current design removed the fabric's only Internet default.\n* Next-hop-self is configured on the edge to spine sessions and operated as designed. Next hop reachability played no role in the fabric failure.\n* No Teraswitch automation run or configuration change triggered the event; the trigger was an external provider maintenance. The vulnerability it exposed, however, was our own: a standing configuration error, the transposed community values, deployed well before the event.\n\n### The default route left on MIA1\n\nThe origin of the route was identified in er100a.mia1's configuration: a static default route left over from site turn-up, still present and still redistributed into BGP.\n\n```\n\"static_routes\": [\n    {\n        \"destination\": \"0.0.0.0/0\",\n        \"nexthop\": \"100.105.0.2\",\n        \"vrf\": \"default\"\n    }\n]\n```\n\nThe static was configured when the site operated on a single transit provider and should have been removed when the second transit was activated. Its next hop, `100.105.0.2`, is a Dallas edge router directly connected to er100a.mia1. Because our edge routers prefer transit-learned defaults over static ones, it sat inactive in the background for months, doing nothing, until the transit provider maintenance withdrew the external default and activated it.\n\nA same-day audit of every edge router found no other static default routes, and edge router configuration templates have been updated to disallow static default routes entirely. IS-IS has also been blocked from redistributing default and static routes.\n\n### How the no-export was applied\n\nRoutes from MIA1's edge routers reached the global reflectors over direct sessions, the temporary arrangement described above. An inbound policy on those sessions stamps arriving routes with regional classification values:\n\n```\nroute-map RM-NA-INBOUND permit 10\n set large-community 20326:301:30 20326:301:20 20326:130:10 additive\n```\n\nThe intended values belong to the `20326:130:x` family, which marks normal routes meant to be sent to all regions:\n\n```\nbgp large-community-list standard REGION-NA-EXPORT permit 20326:130:10\nbgp large-community-list standard REGION-EU-EXPORT permit 20326:130:20\nbgp large-community-list standard REGION-APAC-EXPORT permit 20326:130:30\n```\n\nThe first two values written contain a digit transposition: `20326:301:30` and `20326:301:20` where `20326:130:30` and `20326:130:20` were intended. The third, `20326:130:10`, was written correctly. The `20326:301:x` family is reserved for the opposite purpose, regional export control:\n\n```\nbgp large-community-list standard REGION-NA-NOEXPORT permit 20326:301:10\nbgp large-community-list standard REGION-EU-NOEXPORT permit 20326:301:20\nbgp large-community-list standard REGION-APAC-NOEXPORT permit 20326:301:30\n```\n\nThe global reflectors' export policies toward each region add the well-known `no-export` community to any route carrying that region's value. The goal of this mechanism is containment, not removal: a route marked this way remains present and usable within AS20326 in those regions, but is never exported to transit or peering there. This is a different case from a separate blocking mechanism in the same policies, which prevents a route from being present in a region at all. The edge to spine sessions cross a BGP boundary of the same kind as transit and peering, which is how a flag aimed at external announcements was also able to suppress the fabric signal. The EU policy is shown; the APAC policy is identical in structure:\n\n```\nroute-map RM-EU-OUTBOUND permit 40\n  match large-community REGION-EU-NOEXPORT\n  set community no-export additive\n```\n\nEvery route MIA1 sent over the direct sessions was therefore marked for no-export toward EU and APAC. The NA value, `20326:301:10`, was never written by the transposition, so copies distributed within North America carried no flag. This policy is defined for direct edge to reflector sessions and was attached only at MIA1, the only site operating without its own local reflectors. Validation testing at the time did not cover the edge to global reflector operating mode combined with static default routes existing on the edge router; our testing expansion and configuration sanity checking now do.\n\nNo capture displaying the route's communities survived from the event window. The chain described here was confirmed by replicating the event in a lab environment using pre-outage configuration state.\n\n### The trigger and remaining open items\n\nThe origin and the trigger are now confirmed. The leftover static default on er100a.mia1 activated when the provider maintenance described earlier withdrew the local external default. The route entered BGP, was stamped with the transposed community values on MIA1's direct sessions to the global reflectors, and received the no-export community toward EU and APAC as a result. The transposition was not caught in validation testing because testing did not cover the temporary operating mode of an edge router connected directly to the global reflectors; that testing gap has since been closed. Had MIA1 not been removed from the backbone, the event would eventually have resolved itself: when the provider maintenance completed and the external default returned, it would have displaced the static and the route would have been withdrawn.\n\nThe interaction between the two MIA1 edge routers is resolved: er100b.mia1 never accepted er100a.mia1's static default route, because the BGP policy between the two edge routers intentionally does not share static routes. This is why MIA1 itself initially stayed online: er100b.mia1 continued feeding both MIA1 spines a working default carrying no no-export flag. One item remains under review: why the EU regional reflector at AMS2 emitted the route toward its clients while the second EU regional reflector, hosted in Frankfurt, did not. The AMS2 reflector's running state, which would have answered this directly, was lost when it was powered off during containment. This document will be revised as it closes.\n\n### Contributing factors\n\n* **Readiness signaled with external routes rather than a locally originated default.** Edge routers did not generate an artificial, clean default route of their own to signal readiness to the data center networks. Instead they re-advertised default routes learned from external transit providers, prepared for that role by route-maps. The readiness signal therefore carried external attributes and inherited external failure modes, and a stale or invalid route could still win best path and stand in for a valid signal.\n* **Fail-closed fabric behavior.** With the edges' best default carrying no-export and therefore barred from the fabric sessions, the fabric was left with no default at all and withdrew Internet forwarding entirely. There was nothing to fall back on: routing does not retain a departed route, and no lower-preference default existed to activate in its place. We considered this a critical design flaw and prioritized its correction. The two available solutions were a lower-preference default that activates when the expected one is lost, or force-advertising the default from the edge routers; we chose force-advertisement, deployed as part of the August 12 hardening.\n* **Misleading healthy signals during diagnosis.** L2 transport and private VRF traffic continued to flow from edge to spine, masking that only the `INTERNET` vrf had failed, and management access remained fully functional after failing over to out of band connectivity, suggesting equipment and fabrics were healthy. The failover did raise alarms, but within the wave of critical alarms accompanying the onset, nothing specific to the management network rerouting stood out. Together with a geographic pattern resembling a DDoS event and a recent unrelated IS-IS instability, these signals directed early effort away from fabric Internet routing.\n* **Evidence capture was manual.** Service restoration rightly took priority during the response, and debug capture required manual effort. While our responding engineer guessed correctly that the issue was with MIA1's default export, we do not have a capture of the no-export community on the route from the moments during the outage. We supplemented this by replicating the issue in our lab.\n* **Incorrect generalization from a prior event.** A November 2024 default route incident at FRA2 was caused by our own maintenance tooling in combination with a simultaneous route flap at an outside provider. Its remediation led to the assumption that default route failures would originate from within our own change processes, and the design was never evaluated against a stale or invalid update arriving via the backbone itself. See the following section.\n* **A digit transposition, unprotected by testing.** The inbound policy on the direct edge to global reflector sessions was written with `20326:301:x` values where `20326:130:x` was intended, placing every MIA1 route into the regional no-export control lists. Validation testing did not exercise the temporary operating mode of an edge router connected directly to the global reflectors, so the error was never caught before the event.\n\n## Prior related event and design assumptions\n\nIn the interest of a complete and honest accounting: this was not the first time a default route problem isolated a Teraswitch data center fabric. Teraswitch was aware of a related failure mode from an incident at FRA2 on November 27, 2024, beginning at 09:35 UTC. The conclusions we drew from that event shaped a design assumption that the August 12, 2026 outage proved incorrect. In both events, a default route carrying no-export won selection on a site's edge routers, and the fabric behind them lost its signal. In both, a routine provider event touching the local default exposed the latent flag.\n\n### The November 27, 2024 FRA2 event\n\nThe root cause was a combination of unfortunate provider timing and a logical error in the BGP tooling used to divert traffic away from links about to undergo maintenance:\n\n1. The maintenance tooling set a `no-export` flag \\(BGP community\\) on all routes learned from the AS20326 backbone. Inadvertently, this included the default route \\(`0.0.0.0/0`\\) that is exported toward the data center network to signal that an edge router is ready to pass Internet traffic.\n2. Initially there was no impact: FRA2's edge routers had selected a default route from a local FRA2 Internet transit provider as that signal.\n3. Possibly affected by the same fiber vendor maintenance, one or more local FRA2 transit providers regenerated their default route, causing the FRA2 edge routers to elect a new best default: one generated at a nearby location \\(FRA1\\) and learned over the backbone, still carrying the `no-export` flag.\n4. The FRA2 edge routers therefore stopped exporting a default route toward the FRA2 data center network \\(routers within a different ASN than AS20326\\). The data center concluded that no edge router was available to process Internet traffic, causing a total FRA2 outage until the flag was removed.\n\n### Remediation completed after the 2024 event\n\n* The application of the `no-export` flag to default routes was removed entirely. It has no traffic steering effect, since Internet destinations almost always have a more specific route within the Teraswitch network.\n* A safety was added to the edge routers' route-maps that strips all BGP communities from routes intended to be sent down to the spines, so a control flag can never again suppress the advertisement of the default route unintentionally. As the 2026 event showed, this protection is structurally limited: the well-known no-export community blocks advertisement before any outbound route-map runs, so a route arriving already flagged is never offered to the route-map at all.\n* A route-map on the edge routers was configured to never set the `no-export` or `no-advertise` communities on a default route sent up to the reflectors or to other edge routers.\n\n### The assumption we drew, and why it was wrong\n\nThe 2024 event stemmed from our own tooling, triggered when an outside provider happened to flap its routes at the same time; the fix was within our own change process. From this, Teraswitch concluded that a falsely applied no-export would come from our own edge routers under maintenance, and that tooling and process corrections at the edges were sufficient protection. We never accounted for, or considered, the no-export community being added anywhere else: by a route reflector, or by an external ASN. Nor did we consider a default route arriving as \"local\" but not actually being usable as the readiness signal for the fabric spines. Those assumptions were incorrect. We remained vulnerable to the same flag arriving from any origin other than the one we had specifically resolved.\n\nDuring the August 12, 2026 event the community was already present on the stale MIA1 route when it reached the affected edge routers. It was applied by our own global reflectors' direct-edge import policies. This bypassed all changes and improvements from the 2024 event.\n\nThe global fabric hardening deployed on August 12, 2026 closes this design flaw: the edge routers now force-advertise a locally generated default whenever they are genuinely able to carry Internet traffic, so the fabric always holds a valid default while its edges are healthy, regardless of where the edges learn their own default route or what control communities that route carries.\n\n## Resolution and recovery\n\nEngineers identified the stale route within 10 minutes of onset. MIA1 was removed from the backbone by shutting down the loopback interface used for its SR-MPLS/IS-IS connectivity, and the AMS2 route reflector was powered off, halting further propagation and speeding recalculation. At that moment the responding engineer did not yet know whether MIA1 itself was the source or the AMS2 reflector was malfunctioning; what was known was that the affected sites had learned the route through AMS2's reflection, so both were removed. The team also halted route exports from the global route reflectors in the same window. The only operational change from that halt was peering traffic taking local and regional paths rather than crossing regions; customer and internal routes were maintained throughout by our Continuity Reflectors, which handle only internal and customer routes. With the route's origin now confirmed as er100a.mia1, removing MIA1 from the backbone withdrew the stale default at its source and restored service; powering off the AMS2 reflector removed the EU distribution path and forced recalculation. Reconvergence began at `04:16:15 UTC`. The AMS2 reflector shutdown left a large backlog of route changes, so withdrawing and reprogramming the default route took time, and service returned progressively as each site reconverged on its local default routes. Full recovery, with routes recalculated and traffic patterns back to normal, took 33 minutes from the first alarm. MIA1 returned to the backbone through a maintenance shortly after the event.\n\nLater the same day, working with our network hardware vendor, we deployed a global configuration change across all compute sites, including unaffected data centers, to improve the reliability of the connection between our Internet/backbone edge routers and our data center core networks. With this change in place, a stale or invalid route entering the fabric will no longer prevent traffic forwarding: sites will continue to forward via their local edge routers rather than losing reachability.\n\nThe rollout favored safety over speed. The configuration was updated on the edge routers feeding approximately 62 spine devices globally, one edge router at a time, completing before 15:00 UTC. At that point we did not yet know where the route had come from, so the global route reflectors were kept offline for the duration of the rollout to halt any propagation if it happened again, and the network ran on regional routes only.\n\n### How the fix works\n\nThe previous design had each edge router share the default route it had itself selected as the readiness signal toward the spines. The new design uses the BGP `default-originate` feature on the neighbor group from the edge routers toward the spines. The edge now generates the default route itself, so the route sent into `vrf INTERNET` always carries the local edge router as its next hop and can never inherit a next hop or attributes from a stale or invalid route elsewhere in the backbone.\n\nAdvertisement of this default is controlled by our automation system rather than by a route policy. When maintenance or a traffic drain requires an edge router to stop attracting traffic, the automation deploys a flag on that edge router that stops the advertisement, and removes the flag to restore it. A route-map on the session shapes the attributes of the originated default but does not control whether it is originated. This preserves the original design goal, that an edge router does not attract traffic from the spines before it is ready, retaining the blackhole protection the shared-default design was built for while removing the dependency on a route that transits the backbone.\n\n### Route reflector rebuild and software diversity\n\nAs further hardening, all backbone route reflectors were wiped and fully rebuilt with the latest operating system and reflector software in the week following the event. We also introduced reflector software diversity: a portion of the reflectors now runs a completely different software package, so a defect in any single implementation can no longer affect every reflector at once. This also hedges against the class of defect that lives in an interaction between specific software and firmware versions rather than within any single component.\n\n## Corrective actions\n\n1. Deploy global fabric hardening: replace shared default route signaling with `default-originate` from the edge routers toward the spines, so a stale or invalid route entering the fabric cannot prevent traffic forwarding. \\(**Complete**, Aug 12, 2026\\)\n2. Wipe and fully rebuild all backbone route reflectors with the latest operating system and reflector software, and introduce reflector software diversity by running a portion of the reflectors on a completely different software package. \\(**Complete**, Aug 12, 2026 and following days\\)\n3. Reintegrate MIA1 into the backbone. \\(**Complete**, Post-event maintenance\\)\n4. Correct the transposed community values and audit all route-maps for values colliding with control community families. \\(**Complete**, Week of Aug 10, 2026\\)\n5. Increase logging and automatic debug state capture across backbone devices, including BMP for BGP monitoring, so forensic evidence is preserved without slowing service restoration. \\(**In progress**, Ongoing\\)\n6. Deploy per-VRF default route monitoring: spine and leaf devices raise an alarm when a VRF that previously carried an active default route no longer has one. \\(**In progress**, Ongoing\\)\n7. Move customer DoubleZero interconnects to within the data center fabric, so tunnel sessions no longer depend on paths outside it. \\(**Planning**, TBD\\)\n8. Expand validation testing to cover alternative operating modes, including edge routers connected directly to the global reflectors. \\(**Complete**, Post-event\\)\n9. Audit every edge router for leftover static default routes \\(none besides MIA1 were found\\), and update edge router configuration templates to disallow static default routes entirely. \\(**Complete**, Aug 12, 2026\\)\n10. Block IS-IS from redistributing default routes and static routes. \\(**Complete**, Post-event\\)\n11. Review and rebuild the BGP community scheme network-wide, an ongoing project that predates this event. \\(**In progress**, Ongoing\\)\n12. Reclassify management network failover to out of band connectivity as a higher priority alarm than most other alerts, so it stands out during an alarm storm. \\(**Planning**, TBD\\)\n\n## Lessons learned\n\n* Signal our own gear with a route we originate ourselves. When a default route is used to tell our downstream devices that an edge is ready, that route should be generated locally for that purpose, not borrowed from an external source that carries its own attributes and failure modes.\n* The fate of routing between two devices should be decided between those two devices. A forwarding decision on a local link should not be inherited from a different site, a transit provider, or any other distant route source; the new design enforces this on the edge to spine boundary.\n* Signals that look healthy can hide a broken plane. Traffic on shared links and fully functional management access both masked a failure confined to the `INTERNET` vrf. A failover whose alarms are buried in a larger alarm storm, as happened when the management firewalls moved to out of band connectivity, can easily go unnoticed during the response.\n* Service restoration is priority #1 in these outages; engineers are only worried about restoring access. Evidence preservation must therefore be automatic: increased logging and automatic debug capture will preserve forensic state without slowing recovery.\n* Fixing a root cause is not the same as fixing a failure mode. The 2024 remediation addressed the specific trigger \\(a tooling error\\) rather than the general class of failure \\(any untrustworthy default route reaching the fabric\\). The broader vulnerability remained until August 12, 2026.\n\n## Closing statement\n\nThis event was our mistake, even though triggering it required several small and unrelated conditions to align: a leftover route, a transposed digit, and a routine provider maintenance. We take pride in our choices and designs, and this event has driven us to reevaluate multiple other systems to verify their own integrity. The design carried an assumption about what could break: that any default route would be fully \"cleaned\" by its edge routers before it could affect the fabric behind them. This event proved otherwise: the no-export flag was honored rather than cleaned, withholding the fabric's default while every edge router operated normally. This assumption held for six years, but the 2024 FRA2 event was not properly considered beyond the specific triggers of that day. We are also not blind to what you are thinking after reading this: losing a default route inside our own fabric looks stupid. It is one of the most basic functions in networking, and it was not protected as well as it needed to be. If another provider did this, we would chuckle at it and be glad we had not made such a mistake. We learn from other providers' issues and successes all the time; perhaps someone will learn from ours.\n\nAt the same time, the surrounding design performed the way it was built to. Management access was retained to every device for the full duration, the offending route was identified within ten minutes, and affected sites recovered in roughly half an hour. This network absorbs failures every day, of equipment, routes, fiber optic spans, undersea links, and power, and customers rarely notice, because surviving those failures is where most of the design effort goes. This is the first event of this scale since the backbone took its current form in 2020, in a company operating since 2003.\n\nWe also understand the size and the sensitivity of what runs on this network. We remain unbiased about our customer base and about what customers choose to host with us, but we know exactly what earned that base and what set us apart from our competition: people trust us not to have failures. Simultaneous data center failures, in particular, are not acceptable to us. Our change control and management systems have always been built to make that outcome impossible, and every data center is designed to keep running through the loss of every other data center, in what we call island mode. In this event, several data centers all happily accepted the same bad route, and that has changed our view of how issues can propagate.\n\nWe would have preferred to find this failure mode in validation testing rather than in production, and that is the lasting change: the bar for validation and edge case testing has been greatly raised. The specific condition behind this event can no longer occur, the remaining hardening is being deployed in deliberately sequenced steps so that remediation never becomes its own event, and the network is stronger today than it was on August 11.\n\n**Brendan Mannella**  \nChief Executive Officer\n\n**Nick Zurku**  \nPrincipal Network Architect\n\n## Appendix\n\n### A. Affected sites\n\n* **EU** \\(LON1, AMS1, AMS2, AMS3, DUB1, DUB2, FRA2\\): Internet and inter-site backbone reachability lost\n* **APAC** \\(SGP1, SGP2, TYO1, TYO2, TYO3\\): Internet and inter-site backbone reachability lost\n* **NA** \\(MIA1\\): Temporarily removed from backbone as containment; operated in island mode with Internet routing over local transit only, later reintegrated via maintenance\n* **NA** \\(All other NA sites\\): Not affected\n\n### B. Terminology\n\n* **default route \\(0.0.0.0/0\\):** A catch-all route covering all Internet destinations. Teraswitch uses it internally to signal that an edge router is able to forward traffic to the Internet.\n* **vrf INTERNET:** The routing instance on the spine layer that carries Internet routes for the data center fabric.\n* **route reflector:** A router that redistributes BGP routes between sites so that every router does not need a direct session with every other router.\n* **IS-IS:** The interior routing protocol Teraswitch uses to distribute reachability information inside its own network.\n* **out of band \\(OOB\\):** Management connectivity that is fully independent of the production network, used to reach equipment even when the production network is impaired.\n* **next hop:** The address a router must be able to reach to actually use a route. If the next hop cannot be resolved, the route is invalid and cannot be installed for forwarding.\n* **BGP attributes:** Metadata carried with a route \\(communities, metric, AS path\\) used to identify its origin and control which route is preferred.\n* **no-export:** A standard BGP community instructing routers not to advertise the route across a BGP boundary into another ASN.\n* **edge / spine / leaf:** The three layers of a Teraswitch site: edge routers face the Internet and backbone, spines aggregate the site, and leaf switches connect customer servers.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "critical",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-08-12T18:29:25.451Z",
"resolved_inferred": false,
"started_at": "2026-08-12T05:18:53.835Z",
"state": "postmortem",
"title": "Internet Outage Event at Multiple Sites",
"updated_at": "2026-09-10T20:19:22.400Z",
"url": "https://stspg.io/1rmc0tlrld0b"
},
{
"body": "Teraswitch has completed remediations to remove the bug-triggering configurations from our routers, as well as upgrading our route reflectors and adding diverse vendors on routing software. These stabilizations have proven reliable since July 30th, when we have not seen any additonal invalid packets (IS-IS LSA advertisements) inside our backbone.\n\nWe expect no further impacts from this issue.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-08-18T17:37:54.733Z",
"resolved_inferred": false,
"started_at": "2026-07-29T19:15:40.868Z",
"state": "resolved",
"title": "Global - Internet Instability",
"updated_at": "2026-08-18T17:37:54.747Z",
"url": "https://stspg.io/jwwsk07b22zz"
},
{
"body": "A possible firmware/hardware bug was identified, and Teraswitch NOC elected to do a sequential software upgrade on the redundant switch pair in R102.\n\nAfter the reloads were complete, operations seemed to normalized. We will investigate this network switch pair further to understand the event better.\n\nPlease report any outstanding issues with TYO2 to noc@teraswitch.com or to our team directly.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-28T16:52:29.111Z",
"resolved_inferred": false,
"started_at": "2026-07-28T16:17:12.719Z",
"state": "resolved",
"title": "TYO2 - Network Connectivity Issues with R102 Compute Rack",
"updated_at": "2026-07-28T16:52:29.124Z",
"url": "https://stspg.io/mjnqq334ss3r"
},
{
"body": "Our facility partner has restored power to Teraswitch's SGP2 network racks.\n\nRoot cause: both redundant PDUs in each network rack were fed by power breakout boxes connected to the same overhead electrical bus bar, creating a single point of failure. Facility engineers relocated one breakout box to the correct bus, restoring power to the network equipment and eliminating the single point of failure.\n\nWe expect no further impact at this time. Customer systems did not lose power at any point during this event.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "critical",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-18T01:57:43.885Z",
"resolved_inferred": false,
"started_at": "2026-07-17T23:45:37.835Z",
"state": "resolved",
"title": "SGP2 - Internet Connectivity Loss",
"updated_at": "2026-07-18T01:57:43.902Z",
"url": "https://stspg.io/5z1vg2xwgxvh"
},
{
"body": "Teraswitch monitoring reports that all impacts are currently resolved. Please report any further issues to our team.\n\nAt EWR2, a device sitting between our data center fabric and our internet edge routers experienced a hardware fault that triggered an ASIC-level reload. After the reload, the device came back online in a degraded state and was unable to forward traffic from the data center fabric to the internet edge routers (specifically, it could not process VXLAN-routed traffic).\n\nEWR2 operates six of these core routers. Five were unaffected; however, the impaired unit failed to signal that it was incapable of forwarding traffic. Because downstream routing protocols did not detect this bad state, traffic that landed on the affected device was blackholed.\n\nOur team manually removed the router from the pool of forwarding destinations for this traffic, restoring service. We will perform all necessary follow-up steps to fully resolve the underlying issue and prevent recurrence.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-07-08T18:20:36.465Z",
"resolved_inferred": false,
"started_at": "2026-07-08T17:48:45.507Z",
"state": "resolved",
"title": "EWR2 - Reported Network Connectivity Issues",
"updated_at": "2026-07-08T18:20:36.485Z",
"url": "https://stspg.io/h5nypgf38l9r"
},
{
"body": "The Concerto cable was fully restored - multiple redundant links are now back online.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-06-23T18:43:00.653Z",
"resolved_inferred": false,
"started_at": "2026-06-10T00:24:56.019Z",
"state": "resolved",
"title": "DUB1/DUB2 - Network connectivity issue - backbone loss",
"updated_at": "2026-06-23T18:43:00.678Z",
"url": "https://stspg.io/c1g2q0y75l68"
},
{
"body": "Tokyo to Seattle has been nominal operation for many hours now. We are working with our vendors to understand the issue, but at this time we consider the incident and impact resolved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-06-04T21:11:31.051Z",
"resolved_inferred": false,
"started_at": "2026-06-04T15:27:37.854Z",
"state": "resolved",
"title": "Backbone - Seattle to Tokyo Detour",
"updated_at": "2026-06-04T21:11:31.067Z",
"url": "https://stspg.io/f486k8tttvpb"
},
{
"body": "This incident has been resolved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-05-22T18:22:16.078Z",
"resolved_inferred": false,
"started_at": "2026-05-12T21:18:03.264Z",
"state": "resolved",
"title": "Teraswitch Portal / API - VAN1 management temporarily unavailable",
"updated_at": "2026-05-22T18:22:16.096Z",
"url": "https://stspg.io/pgwmg9vtzhzp"
},
{
"body": "This incident has been resolved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-05-22T18:24:53.128Z",
"resolved_inferred": false,
"started_at": "2026-05-07T02:03:18.963Z",
"state": "resolved",
"title": "TYO2 - Core Router Reload",
"updated_at": "2026-05-22T18:24:53.144Z",
"url": "https://stspg.io/kqhs7qfkzbw2"
},
{
"body": "This incident has been resolved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-05-01T17:30:06.098Z",
"resolved_inferred": false,
"started_at": "2026-04-30T23:33:00.510Z",
"state": "resolved",
"title": "TYO1 - Network Connectivity Issue - Single Rack",
"updated_at": "2026-05-01T17:30:06.111Z",
"url": "https://stspg.io/gq1cw26rxxsr"
},
{
"body": "# Root Cause Analysis: AMS1 Connectivity Outage\n\n**Incident Date:** March 26, 2025 **Duration:** ~1h 21m \\(22:08 UTC \u2013 00:14 UTC \\+1\\) **Severity:** Critical \u2013 Full site connectivity loss **Affected Site:** AMS1 \u2013 Amsterdam, Netherlands **Status:** Resolved\n\n## Executive Summary\n\nOn March 26, 2025, Teraswitch's AMS1 facility experienced a complete loss of backbone connectivity lasting approximately 1 hour and 21 minutes. The outage was caused by a scheduled fiber vendor maintenance window that simultaneously impacted both the primary and what was believed to be a diverse redundant fiber path between AMS1 and the rest of the backbone. Investigation revealed that due to a documentation and handoff error at the fiber vendor dating back over a year, the AMS1\u2013AMS2 fiber span had never been migrated to the intended diverse path as part of the AMS3 ring buildout. As a result, both affected spans shared physical infrastructure, eliminating the redundancy intended to protect against exactly this type of event.\n\nConnectivity was restored within the maintenance window after a rapid joint audit with the fiber vendor confirmed the provisioning discrepancy, and the AMS1\u2013AMS2 span was moved to its correct diverse path.\n\n## Background\n\nTeraswitch's Amsterdam backbone originally consisted of a single dark fiber span between AMS1 and AMS2. When AMS3 was later brought online, the network design called for a three-node fiber ring with fully diverse physical paths between all sites to provide redundant backbone connectivity.\n\nTo support this design, the fiber vendor was engaged to:\n\n1. Provision a new AMS1\u2013AMS3 span\n2. Provision a new AMS2\u2013AMS3 span\n3. Reroute the existing AMS1\u2013AMS2 span onto a physically diverse path\n\nDue to an internal handoff and documentation error within the fiber vendor, step 3 was not completed. The AMS1\u2013AMS2 span remained on its original physical route. Teraswitch was not made aware of this omission, and the span continued to operate normally for over a year. Because it carried live traffic and appeared correctly in topology, it was not identified as incorrectly provisioned during subsequent audits.\n\n## Timeline of Events\n\n| Time \\(UTC\\) | Event |\n| --- | --- |\n| Prior to March 26 | Fiber vendor schedules routine maintenance affecting Amsterdam infrastructure |\n| ~22:08 | AMS1 backbone connectivity lost. Both AMS1\u2013AMS2 and AMS1\u2013AMS3 paths go down simultaneously. Teraswitch NOC begins triage. |\n| ~22:53 | Fiber vendor engaged and confirms maintenance is impacting both spans. Root cause identified as shared physical infrastructure due to the original provisioning error. |\n| ~00:14 \\+1 | Fiber vendor migrates AMS1\u2013AMS2 span to the correct diverse physical path. Connectivity restored. Monitoring confirmed stable. |\n\n## Root Cause\n\n**Primary cause:** An internal documentation and handoff failure at the fiber vendor resulted in the AMS1\u2013AMS2 dark fiber span never being migrated to its intended physically diverse route during the AMS3 ring buildout. Both the AMS1\u2013AMS2 and AMS1\u2013AMS3 spans shared common physical infrastructure, making the designed ring topology's redundancy ineffective.\n\n**Contributing factor:** Because the span was operationally active and traffic was flowing normally, the provisioning error went undetected across both Teraswitch and vendor records for over a year.\n\nWhen the scheduled maintenance affected the shared physical infrastructure, both paths were impacted simultaneously, leaving AMS1 with no available backbone connectivity.\n\n## Impact\n\n* **AMS1 customers** experienced a complete loss of inbound and outbound connectivity for approximately 1 hour and 21 minutes.\n* No data loss or hardware damage occurred.\n* All other Teraswitch sites were unaffected.\n\n## Resolution\n\nWorking jointly with the fiber vendor during the incident, Teraswitch engineers and vendor technicians audited the physical path assignments for all Amsterdam spans. The discrepancy between the intended and actual routing of the AMS1\u2013AMS2 span was identified. The vendor migrated the span to the correct physically diverse path, restoring independent redundant connectivity across the AMS1\u2013AMS2\u2013AMS3 ring as originally designed.\n\n## Corrective Actions\n\n| Action | Owner | Status |\n| --- | --- | --- |\n| Confirm and document physical path diversity for all three Amsterdam spans with fiber vendor | Teraswitch / Fiber Vendor | Complete |\n| Obtain updated as-built fiber records from vendor reflecting correct path assignments | Fiber Vendor | Complete |\n| Audit all other Teraswitch sites for similar provisioning discrepancies against vendor records | Teraswitch | Complete |\n| Establish a fiber path verification checklist for all future vendor provisioning work prior to accepting new spans | Teraswitch | Planned |\n| Add physical diversity validation to change management process for any future ring or redundancy buildouts | Teraswitch | Planned |\n| Incorporate optical span latency validation against fiber path build sheets as an acceptance criterion for new span provisioning | Teraswitch | Planned |\n\n## Lessons Learned\n\n* **Operational traffic is not proof of correct provisioning.** A span can carry live traffic for an extended period while still being routed incorrectly relative to its intended physical diversity design.\n* **Redundancy assumptions must be periodically verified against vendor as-built records**, not solely inferred from operational status.\n* **Fiber vendor handoffs require explicit acceptance criteria** including documented physical path confirmation before provisioning work is considered complete.\n* **Span latency is a low-cost signal for path verification.** In post-incident review, Teraswitch noted that the measured propagation latency on the AMS1\u2013AMS2 span was slightly lower than expected based on the fiber path build sheet for the intended diverse route. This discrepancy, while subtle, was consistent with the span still traversing the shorter original path. Validating measured latency against estimated values from build sheets at provisioning acceptance could have surfaced this error significantly earlier. This check will be incorporated into the span acceptance process going forward.\n\n_RCA prepared by Teraswitch Network Engineering. For questions contact the NOC or network architecture team._",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "critical",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-04-02T22:38:06.946Z",
"resolved_inferred": false,
"started_at": "2026-03-26T22:08:43.581Z",
"state": "postmortem",
"title": "AMS1 - Loss of Connectivity",
"updated_at": "2026-04-02T22:38:47.029Z",
"url": "https://stspg.io/w5r5d85ryn3m"
},
{
"body": "This incident has been resolved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "major",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-03-25T22:00:05.297Z",
"resolved_inferred": false,
"started_at": "2026-03-19T21:43:43.000Z",
"state": "resolved",
"title": "EWR2 - Network Connectivity Issues",
"updated_at": "2026-03-25T22:00:05.314Z",
"url": "https://stspg.io/24zh1n01w7my"
},
{
"body": "This incident has been resolved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-03-25T22:03:03.321Z",
"resolved_inferred": false,
"started_at": "2026-03-17T20:05:09.047Z",
"state": "resolved",
"title": "PIT1 - Sporadic Internet Connectivity Issues",
"updated_at": "2026-03-25T22:03:03.335Z",
"url": "https://stspg.io/6ct1khtkhm12"
},
{
"body": "Cogent has recovered as of 4:33am Eastern. Services should be normalized at this time.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-02-20T11:32:19.000Z",
"resolved_inferred": false,
"started_at": "2026-02-20T07:47:15.436Z",
"state": "resolved",
"title": "VAN1 - Connectivity Issues",
"updated_at": "2026-02-20T11:32:38.165Z",
"url": "https://stspg.io/9njyx1jvk9ck"
},
{
"body": "This incident has been resolved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-02-24T15:44:30.306Z",
"resolved_inferred": false,
"started_at": "2026-02-20T07:01:09.407Z",
"state": "resolved",
"title": "APAC - Loss of Tokyo to Seattle Backbone Paths",
"updated_at": "2026-02-24T15:44:30.323Z",
"url": "https://stspg.io/7pv903mcl15t"
},
{
"body": "We have successfully brought online a second diverse sub-sea cable path and latency has returned to normal. Once the original severed cable comes back online, we will now have full path redundancy.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-03-09T02:08:41.160Z",
"resolved_inferred": false,
"started_at": "2026-02-19T02:08:01.426Z",
"state": "resolved",
"title": "Subsea Cable Fault - Singapore to Europe",
"updated_at": "2026-03-09T02:08:41.182Z",
"url": "https://stspg.io/ypljm16m7tcq"
},
{
"body": "This incident has been resolved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-02-05T16:46:00.000Z",
"resolved_inferred": false,
"started_at": "2026-02-04T10:45:55.000Z",
"state": "resolved",
"title": "Multiple US Backbone Connection Losses",
"updated_at": "2026-02-09T16:46:37.827Z",
"url": "https://stspg.io/7kdbbsz7xz98"
},
{
"body": "Cloudflare has resolved the issue, and console.tsw.io is accessible once again as of 22:28 UTC.  This incident has been resolved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "critical",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-02-03T22:37:52.956Z",
"resolved_inferred": false,
"started_at": "2026-02-03T19:54:44.362Z",
"state": "resolved",
"title": "Teraswitch Console (console.tsw.io) - Inaccessible",
"updated_at": "2026-02-03T22:37:52.974Z",
"url": "https://stspg.io/6f0gg7pnq1s3"
},
{
"body": "This incident has been resolved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-01-30T10:30:32.476Z",
"resolved_inferred": false,
"started_at": "2026-01-30T09:53:25.098Z",
"state": "resolved",
"title": "West Coast Packet Loss",
"updated_at": "2026-01-30T10:30:32.497Z",
"url": "https://stspg.io/jfc4hn87y9mn"
},
{
"body": "At 2PM UTC, Teraswitch reloaded software related to our route servers to fix \"stuck\" transport tunneling sessions. This may have caused interruptions to ULL/HFT and transport services over our network. \n\nWe believe the first stuck session started approximately 6 hours before, with more dropping off over the next few hours.\n\nAs of this moment, all stuck sessions appear to be resolved, and traffic is passing normally and via ULL links.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-01-28T14:17:19.376Z",
"resolved_inferred": false,
"started_at": "2026-01-28T14:17:19.327Z",
"state": "resolved",
"title": "Transport/Ultra Low Latency Services Impacted",
"updated_at": "2026-01-28T14:17:19.385Z",
"url": "https://stspg.io/fb748gs1wxtq"
},
{
"body": "This incident has been resolved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-01-26T13:56:03.410Z",
"resolved_inferred": false,
"started_at": "2025-12-28T02:19:14.019Z",
"state": "resolved",
"title": "SGP1/SGP2 - Backbone Isolation Event",
"updated_at": "2026-01-26T13:56:03.425Z",
"url": "https://stspg.io/fmq26pds5bg4"
},
{
"body": "The cause of the increased errors was identified and fixes put in place.  This issue is resolved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-12-05T23:46:17.714Z",
"resolved_inferred": false,
"started_at": "2025-12-05T21:53:33.000Z",
"state": "resolved",
"title": "Teraswitch Console - Increased Errors",
"updated_at": "2025-12-05T23:46:17.732Z",
"url": "https://stspg.io/ccc5bmh6hk65"
},
{
"body": "Fault has cleared and cable returned to service.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-11-29T13:02:22.422Z",
"resolved_inferred": false,
"started_at": "2025-11-29T05:21:58.947Z",
"state": "resolved",
"title": "Seattle -> Tokyo Cable Fault",
"updated_at": "2025-11-29T13:02:22.441Z",
"url": "https://stspg.io/gd6z69yw6bhv"
},
{
"body": "Cloudflare updated their incident report that their services are now operating normally - they will post a final update to the incident once their investigation has concluded.  Teraswitch monitoring confirms our console / API services have stabilized with no further errors.  This incident is now resolved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "critical",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-11-18T18:09:11.691Z",
"resolved_inferred": false,
"started_at": "2025-11-18T12:08:44.233Z",
"state": "resolved",
"title": "Teraswitch Console & API - Increased Errors / Unavailable",
"updated_at": "2025-11-18T18:09:11.705Z",
"url": "https://stspg.io/p6rgrsl2g0zk"
},
{
"body": "This incident is now considered resolved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-11-17T13:45:35.069Z",
"resolved_inferred": false,
"started_at": "2025-11-14T15:17:21.870Z",
"state": "resolved",
"title": "Internet Connectivity Instability",
"updated_at": "2025-11-17T13:45:35.085Z",
"url": "https://stspg.io/bs1glxq85nkc"
},
{
"body": "This incident has been resolved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "critical",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-10-29T20:27:17.331Z",
"resolved_inferred": false,
"started_at": "2025-10-29T16:13:07.000Z",
"state": "resolved",
"title": "Teraswitch Console (console.tsw.io) - Inaccessible",
"updated_at": "2025-10-29T20:27:17.349Z",
"url": "https://stspg.io/r3864x53553k"
},
{
"body": "This detour has already cleared and was a momentary disruption. TeraSwitch will engage our transport vendor for a better understanding of this incident.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-10-29T00:36:51.952Z",
"resolved_inferred": false,
"started_at": "2025-10-29T00:20:29.380Z",
"state": "resolved",
"title": "New York <-> London Detoured",
"updated_at": "2025-10-29T00:36:51.966Z",
"url": "https://stspg.io/s99b33qd8pr4"
},
{
"body": "All operations of Console services have been restored. Users should not see any issues managing their accounts and services.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "critical",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-10-01T23:30:35.118Z",
"resolved_inferred": false,
"started_at": "2025-10-01T21:52:17.844Z",
"state": "resolved",
"title": "Teraswitch Console / console.tsw.io - Inaccessible",
"updated_at": "2025-10-01T23:30:35.134Z",
"url": "https://stspg.io/c5sh8tqfjmtg"
},
{
"body": "This temperature issue was quickly resolved yesterday and was identified to be caused by maintenance work being completed on a nearby HVAC unit.\n\nThe temperature increased due to lower airflow. At this time, all work is complete and the HVAC maintenance is completed. Our facility vendor is working to increase airflow so that in the future, specific units aren't required to supply the proper airflow.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-10-01T14:13:26.134Z",
"resolved_inferred": false,
"started_at": "2025-09-30T19:01:50.246Z",
"state": "resolved",
"title": "SLC1 - Elevated Temperatures in Datahall",
"updated_at": "2025-10-01T14:13:26.151Z",
"url": "https://stspg.io/yfxkglb87s4p"
},
{
"body": "An alternative undersea cable has been established to restore this route.\n\nSGP<->TYO connectivity is now via backbone.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2026-01-26T13:57:12.749Z",
"resolved_inferred": false,
"started_at": "2025-09-26T00:00:06.000Z",
"state": "resolved",
"title": "Singapore <-> Tokyo Detoured",
"updated_at": "2026-01-26T13:57:12.763Z",
"url": "https://stspg.io/krwm8qjqqbnb"
},
{
"body": "The path has been restored and operations are normal between NJ and Ashburn.\n\nWe are awaiting the provider's response about the cause of the disturbance.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-09-24T20:10:05.415Z",
"resolved_inferred": false,
"started_at": "2025-09-24T19:58:34.549Z",
"state": "resolved",
"title": "New Jersey <-> Ashburn Detoured",
"updated_at": "2025-09-24T20:10:05.434Z",
"url": "https://stspg.io/3fpv56f21wp1"
},
{
"body": "The fiber cut has been restored and routing is now normalized.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-09-05T00:59:30.114Z",
"resolved_inferred": false,
"started_at": "2025-09-04T21:43:16.914Z",
"state": "resolved",
"title": "Amsterdam <-> Frankfurt Detoured",
"updated_at": "2025-09-05T00:59:30.128Z",
"url": "https://stspg.io/rgpb8b21wtbl"
},
{
"body": "Our network transport provider was able to roll back their changes with no impact. At a later date they will replace the malfunctioning hardware.\n\nTeraswitch has normalized operations in Frankfurt.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-09-02T15:26:32.070Z",
"resolved_inferred": false,
"started_at": "2025-09-01T23:18:50.211Z",
"state": "resolved",
"title": "Frankfurt - Internet Traffic Diversion",
"updated_at": "2025-09-02T15:26:32.093Z",
"url": "https://stspg.io/kddnbhmg1gxh"
},
{
"body": "Undersea cable operations have been restored, and traffic is stable and normalized along this path.\n\nTokyo to Seattle is restored.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-10-17T18:38:29.290Z",
"resolved_inferred": false,
"started_at": "2025-09-01T05:15:25.431Z",
"state": "resolved",
"title": "Seattle <-> Tokyo Detoured",
"updated_at": "2025-10-17T18:38:29.307Z",
"url": "https://stspg.io/z58l5fm2rg0f"
},
{
"body": "This incident has been resolved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-08-26T18:30:13.008Z",
"resolved_inferred": false,
"started_at": "2025-08-26T14:22:32.655Z",
"state": "resolved",
"title": "SGP1 - Internet Traffic Latency",
"updated_at": "2025-08-26T18:30:13.026Z",
"url": "https://stspg.io/s9hkrx5vdxzv"
},
{
"body": "This incident has been resolved.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "minor",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-09-02T15:26:56.355Z",
"resolved_inferred": false,
"started_at": "2025-08-25T22:13:40.000Z",
"state": "resolved",
"title": "Multiple Sites - Network Instability Issues",
"updated_at": "2025-09-02T15:26:56.377Z",
"url": "https://stspg.io/gmb2j2z9g6b0"
},
{
"body": "The fiber has been restored.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-08-20T08:21:20.511Z",
"resolved_inferred": false,
"started_at": "2025-08-19T23:00:57.000Z",
"state": "resolved",
"title": "Dublin <-> Amsterdam Intra-Market Connectivity Detoured",
"updated_at": "2025-08-20T08:21:20.528Z",
"url": "https://stspg.io/gz6g055lpwhh"
},
{
"body": "The fiber fault has been cleared and traffic is normalized.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-08-19T18:03:36.134Z",
"resolved_inferred": false,
"started_at": "2025-08-19T08:45:26.000Z",
"state": "resolved",
"title": "Frankfurt <-> Amsterdam Intra-Market Connectivity Detoured",
"updated_at": "2025-08-19T18:03:36.150Z",
"url": "https://stspg.io/pm0bnnvzmrdk"
},
{
"body": "The subsea cable repair is done, the connections have been stable for nearly 24 hours, and traffic has been normalized across the link.\n\nThis concludes this event.",
"first_seen": "2026-09-04T07:06:16Z",
"impact": "none",
"last_seen": "2026-09-16T12:28:20Z",
"resolved_at": "2025-08-20T18:11:40.972Z",
"resolved_inferred": false,
"started_at": "2025-08-12T20:19:32.144Z",
"state": "resolved",
"title": "New York <-> London Detoured",
"updated_at": "2025-08-20T18:11:40.990Z",
"url": "https://stspg.io/mjsm35s1w919"
}
]
}