Summary

  • ThousandEyes reported two distinct Comcast outages on 8 and 9 November 2021. The first began at approximately 9:44 p.m. Pacific time on 8 November and ended at approximately 10:48 p.m. The second began at approximately 5:05 a.m. on 9 November and ended at approximately 6:15 a.m.[1]

  • During the first event, external tests showed packet loss on paths traversing Comcast's Sunnyvale core. Some traffic that initially used other paths remained successful, then failed after it was rerouted through Sunnyvale. That sequence is evidence about observed forwarding paths, not a complete internal Comcast topology or configuration record.[1]

  • The second event had a broader observed footprint. ThousandEyes reported that some Central and Eastern U.S. traffic was temporarily directed toward Sunnyvale even when the endpoints were far from California. Some paths alternated between complete loss and successful reachability, behavior the analysis described as possibly associated with control-plane churn.[1][3]

  • A later ThousandEyes review attributed the incident to an inadvertently exceeded routing-table limit.[2] The public packet does not identify the exact device, table, configured threshold, software behavior, command, change owner, vendor or approval sequence. The limit explanation must therefore remain attributed rather than presented as a complete Comcast postmortem.

  • A routing-table limit is an accountability control because operators can measure table occupancy, growth rate, reserved headroom, alarm thresholds, failure behavior and recovery. Whether those measurements existed or were effective inside Comcast is not disclosed by the frozen sources.

  • Rerouting did not always produce an independent path. Observed traffic that was redirected into the impaired Sunnyvale core also failed. Continuity therefore depends on failure-domain separation, not merely on the existence of another route calculation.[1]

  • ARIN's RDAP record for AS7922 supplies network-resource attribution context.[9] It does not reveal the live routes installed in a Comcast router, internal route-reflector state, topology, table utilization or the forwarding outcome of a particular packet.

  • IETF documents explain BGP operation, convergence, route reflection, graceful change and failure detection.[10]-[20] They supply control vocabulary and later or general design context. They do not prove which mechanisms Comcast deployed in November 2021, and they are not retroactive findings of fault.

  • FCC outage-reporting rules establish an accountability record for qualifying communications outages.[7][8] The public sources used here do not disclose Comcast's confidential filing for this event. A reporting obligation cannot be treated as a public technical postmortem.

  • The evidence supports a measurable operational conclusion, not an accusation. Comcast controlled its internal routing capacity, topology, alarms, change process, customer communication and recovery. External observers controlled their measurement methods. Regulators controlled reporting requirements. The available record does not establish intent, negligence, legal liability or individual responsibility.

The accountability question

The November 2021 incidents matter because the Internet did not fail as an abstract cloud. Specific traffic paths stopped forwarding, some alternate calculations directed traffic into the same troubled core, and service returned after the network state changed. That is an observable operational sequence. It creates a narrower and more useful accountability question than asking whether a large provider should ever experience an outage.

The question is whether the controls that governed routing-table capacity and failure-domain behavior were testable before the event. A provider can know how many routes a device or process supports. It can know current occupancy, the rate at which state grows, the amount of capacity reserved for convergence, the behavior at warning and hard limits, and whether redundant control elements share the same ceiling. It can test whether a failed node or region causes traffic to enter a genuinely independent path. It can preserve evidence showing when alarms fired, who acted, what changed and how forwarding recovered.

Public evidence does not reveal Comcast's answers to those questions. ThousandEyes supplied external observations from its own vantage points and later attributed the incident to a routing-table limit.[1][2] It did not publish Comcast's configuration archive, internal telemetry, approval records or full incident review. Comcast's architecture pages describe a large and distributed network, but they were not written as the postmortem for these two events.[5][6] Accountability analysis must therefore separate what can be observed from what remains inside the operator's evidence boundary.

That separation is not a reason to abandon analysis. It defines the correct unit of responsibility. Comcast controlled the internal system in which the limit was reached and the recovery was executed. ThousandEyes controlled its measurements and interpretation, not Comcast's routers. ARIN controlled the accuracy and availability of the AS7922 registration record, not the routes installed in the Sunnyvale core.[9] The FCC controlled reporting requirements and access rules, not live forwarding decisions.[7][8]

The resulting thesis is practical: a continuity claim is credible only when capacity, topology and recovery can be demonstrated in running systems. A design diagram may show multiple nodes. A routing protocol may calculate multiple paths. A registry may accurately identify an autonomous system. None of those facts by itself proves that traffic will avoid a common impaired domain when a table limit is reached.

A two-event forensic timeline

The chronology comes from external measurement and must remain labeled that way. A path-monitoring platform sees selected tests from selected vantage points. It can reveal packet loss, path changes and recurring patterns. It cannot see every internal command, every route in every table or every customer session. The following timeline records observations without converting them into an imagined internal log.

Time Evidentiary event
Before 9:44 p.m. Pacific, 8 November The public packet does not identify an initiating change, table-growth event, device alarm or internal maintenance action. External paths used by the later analysis were functioning before the observed loss.
Approximately 9:44 p.m. ThousandEyes placed the beginning of the first outage at about this time. Tests whose traffic traversed the Sunnyvale core began to show loss.[1]
Approximately 9:44-9:46 p.m. Some neighboring paths outside Sunnyvale continued to work. The observation matters because it shows that the failure was not initially uniform across every measured path.[1]
From approximately 9:46 p.m. Some traffic was rerouted through Sunnyvale and then experienced the same total packet loss. The external record shows a path change followed by failure, but not the internal decision or route that caused each change.[1]
Approximately 10:48 p.m. The first observed outage ended. Previously rerouted paths returned to earlier routes, while some Sunnyvale-traversing traffic used a different set of Sunnyvale nodes. The public record does not identify the corrective command or exact convergence sequence.[1]
Between the events The packet does not disclose whether a shared condition persisted, whether a change was attempted, or whether the second event had the same immediate trigger. The two incidents showed similar behavior, but similarity is not proof of one uninterrupted internal cause.
Approximately 5:05 a.m., 9 November The second outage began. ThousandEyes observed complete loss on some paths traversing Sunnyvale.[1]
During the second event Some traffic from other U.S. regions was redirected toward Sunnyvale and failed. Chicago-to-Chicago traffic was among the examples used to illustrate the unexpected geographic path.[1]
During the second event Some measured paths alternated between loss and successful reachability. ThousandEyes discussed control-plane churn as a possible explanation for this changing behavior.[1]
Approximately 6:15 a.m. The second observed outage ended and impacted paths again reached their destinations. The public packet does not disclose whether recovery resulted from a rollback, capacity change, process restart, route withdrawal or another action.[1]
Later review ThousandEyes' annual review attributed the incident to an inadvertently exceeded routing-table limit.[2] This later explanation supplies the frozen mechanism but not a full internal root-cause tree.

This timeline supports three conclusions. First, the two events were separated in time and should not be collapsed into one continuous outage without internal evidence. Second, the Sunnyvale core was central to the observed failures. Third, rerouting could increase the affected set when the alternate calculation sent traffic into the same impaired core.

It does not support a claim that every Comcast subscriber was offline, that every path traversed Sunnyvale, or that the limit affected every router in the same way. It does not reveal whether a hard table ceiling caused routes to be rejected, withdrawn, flushed or repeatedly recalculated. It does not establish whether the table was a BGP routing information base, a forwarding table, a platform-specific structure or another control-plane resource. Those distinctions remain essential unknowns.

What external path evidence can establish

External measurements are strongest when they describe forwarding outcomes. A test sends traffic from a known vantage point toward a known destination, records hops and loss, and compares the path before, during and after an incident. When many tests share an affected interface or location, the evidence can identify a common observable failure point. ThousandEyes describes this method as aggregating measurements across its platform to detect traffic and routing outages.[4]

For Comcast's first event, the comparison between successful paths outside Sunnyvale and failing paths through Sunnyvale creates a meaningful boundary. It indicates that the observed failure followed path placement. When some neighboring traffic was later redirected through Sunnyvale and then failed, the sequence showed that the alternate route did not escape the affected domain.[1]

The second event added a geographic anomaly. Some traffic with endpoints in the central or eastern United States was observed traversing Sunnyvale. A path can be technically valid while being operationally undesirable during a failure. BGP and internal routing systems choose paths under configured policy and available state; they do not understand a customer's intuitive expectation that local traffic should remain geographically local.[10] The accountability control is therefore not intuition. It is a testable policy and topology requirement.

External path evidence has limits. A visible hop may not answer every question about encapsulation, internal labels, route reflection or equal-cost multipath. Nonresponding interfaces can complicate interpretation. A path seen by one vantage point is not a universal route. Packet loss at or after a named hop does not always prove that the responding interface caused the loss. ThousandEyes' conclusions should be read as measurement-provider observations, not privileged access to every router.

Those limits make corroboration and retention important. Operators can preserve their own route state, interface counters, table occupancy and change records alongside independent path data. A later review can then test whether an external path shift corresponds to a known internal event. Without that joined evidence, outsiders can identify the failure domain but not reconstruct the full control sequence.

Routing-table capacity is a continuity control

Routing systems store several kinds of state. BGP speakers receive updates, apply policy, select paths and advertise permitted results.[10] Implementations may maintain received routes, accepted routes, selected routes and forwarding entries in separate structures. A route reflector can reduce the need for a full mesh of internal BGP sessions, while also becoming part of the distribution path for routing information.[12] Hardware and software impose limits on memory, forwarding entries, process resources and supported routes.

The phrase "routing-table limit" therefore needs precision. A configured maximum can be a deliberate guardrail. A platform capacity can be a hard engineering boundary. A process can exhaust memory before a nominal route count is reached. A session maximum-prefix feature can shut a session or warn an operator. A forwarding table can have different capacity from a control-plane table. The frozen public evidence does not say which condition occurred in Comcast's network.

That uncertainty does not make capacity unauditable. An operator can record the identity of each relevant table, its supported and configured limits, normal occupancy, peak occupancy, reserved convergence margin and expected growth. It can define warning thresholds below the failure point and test the alert path. It can simulate a controlled increase in route state and verify whether the device rejects only the excess, protects established forwarding, restarts a process, withdraws routes or propagates churn.

Headroom should be defined against a failure scenario, not an average day. During convergence, a router may temporarily retain old and new paths. Maintenance may cause alternate routes to appear. A policy error may increase accepted state. A route-reflector change may alter which paths are visible. The correct margin therefore includes transient state and the time required for an operator to act.

A useful capacity record would contain at least six measurements:

  1. current occupancy for each routing and forwarding structure;
  2. the hard platform limit and any lower configured limit;
  3. the highest transient occupancy observed during tested convergence;
  4. the warning threshold and verified alert-delivery time;
  5. the documented behavior at the warning and hard limits; and
  6. the recovery procedure, including the evidence required before traffic is returned.

The November incident makes this record consequential because the observed failure did not remain local to traffic already using Sunnyvale. Some traffic was redirected into the core and failed there.[1] If a table limit in one node or cluster can attract additional paths during convergence, the limit becomes a blast-radius control. Capacity testing must ask not only whether one device survives, but also how the rest of the network reacts to its partial failure.

No source in the packet establishes that Comcast lacked these measurements. The evidence shows that an exceeded limit was later cited and that externally observed paths failed. The accountability finding is that the relevant measurements should be reviewable. It is not that their absence has been proven.

Why rerouting was not independent resilience

Networks are commonly described as resilient because traffic can take another path. That statement omits the most important question: independent from what? Two paths can use different interfaces while sharing a route reflector, software release, table ceiling, power domain, metro core, maintenance process or configuration source. A new path calculation can therefore preserve the same underlying failure.

The first Comcast event offers a concrete example. Some traffic outside Sunnyvale initially remained successful. After it was redirected through Sunnyvale, it also failed.[1] The protocol found a route, but the route entered an impaired domain. From the user's perspective, the existence of a second calculation did not create continuity.

Failure-domain analysis should be performed at several layers. Physical diversity asks whether links, sites and power systems are separate. Control-plane diversity asks whether route distribution and decision processes can fail independently. Capacity diversity asks whether alternate nodes have independent and sufficient table headroom. Operational diversity asks whether one change, automation system or approval can affect all supposed alternatives. Observability diversity asks whether monitoring remains available when the production control plane is impaired.

A Clos-style core can provide multiple paths and horizontal scale. ThousandEyes discussed Comcast's use of a spine-leaf design when interpreting the event.[1] That architectural context explains why node-level and fabric-level behavior matter. It does not disclose the exact production topology of the affected core or prove that every path shared one control dependency.

The right verification is adversarial but bounded. Operators can remove a node, isolate a route reflector, constrain a table, delay an update and observe where traffic moves. The test should confirm that the alternate path avoids the original physical and logical failure domain, has adequate capacity, and does not create an unexpected geographic detour. Results should be captured in forwarding-plane measurements from inside and outside the network.

Resilience is demonstrated when the alternate path carries traffic under the tested fault. It is not demonstrated when a topology diagram contains multiple lines.

Convergence, route reflection and changing paths

BGP does not update the entire Internet or a large internal network instantaneously. Routers receive changes at different times, apply local policy and advertise new results. RFC 4277 surveys convergence behavior and the delays or transient states that can occur after routing changes.[11] Route reflectors change the propagation structure inside an autonomous system by allowing clients to exchange routes without a full internal mesh.[12]

ThousandEyes observed some Comcast paths alternating between complete loss and normal reachability during the second event and identified control-plane churn as a possible explanation.[1] The public evidence does not show the precise updates responsible. It nevertheless establishes why convergence behavior belongs in a capacity review. A system near a limit may react differently as old and new paths coexist or as sessions reset and repopulate state.

Graceful-restart and graceful-shutdown mechanisms address particular transition problems. RFC 4724 describes preserving forwarding state during certain BGP restarts.[13] RFC 6198 sets requirements for reducing traffic loss when a BGP session is intentionally shut down, and RFC 8326 specifies a graceful-shutdown mechanism.[15][19] These documents do not prove that the mechanisms were relevant, available or deployed in Comcast's incident.

They do show that "the protocol converged" is not a complete operational standard. A convergence event can involve packet loss, transient loops, stale state or path changes that violate intended locality. A tested network should define acceptable convergence time and loss for each failure class. It should also define what happens when a control-plane resource, rather than a link, reaches a limit.

Route reflection deserves specific evidence because logical redundancy can still share distribution state. An operator should know which clients depend on each reflector, whether alternate reflectors have independent capacity, how paths are selected when one reflector loses state, and how a limit alarm changes propagation. The article does not assert that a route reflector caused Comcast's incident. It identifies the type of dependency that a routing-table-limit postmortem should examine.

Detection must survive the failure it reports

Fast detection is only useful when the alert reaches an operator and identifies the affected control surface. Bidirectional Forwarding Detection can provide rapid detection of certain forwarding-path failures.[14] It does not diagnose a routing-table limit by itself. Device telemetry can report table occupancy and process health. Route collectors and external path tests can reveal changes in reachability. Customer reports can show service symptoms. Each source sees a different part of the event.

An accountable monitoring design connects these layers. A table warning should identify the device, structure, current value, configured limit and trend. A routing alarm should show what state changed. A path alarm should show which destinations and regions lost reachability. A customer-impact system should connect the network event to affected services without claiming more users than the evidence supports.

Monitoring also needs an independent delivery path. If alerts, dashboards, authentication or incident chat depend on the same impaired network, the operator may lose the tools needed to recover. The public Comcast packet does not say whether that happened. It remains a measurable continuity requirement derived from the failure class, not an allegation about the event.

External measurements contribute a separate reality check. ThousandEyes could compare successful and failing paths before Comcast published a detailed explanation.[1][4] An operator can use similar external evidence to test whether an internal "green" status corresponds to successful forwarding. The final recovery gate should require both internal stability and external reachability from multiple regions.

That gate prevents a common closeout error: declaring recovery when a control process has restarted but forwarding remains unstable. In the Comcast timeline, the observable end state was that affected paths again reached their destinations.[1] The public record does not reveal Comcast's internal declaration criteria, so no comparison can be made. The incident still illustrates why forwarding evidence belongs in the criteria.

Registry evidence and running reality

ARIN's RDAP service records AS7922 as a registered autonomous system resource.[9] Such records matter. They help operators and investigators identify the organization associated with a number resource, maintain contacts and distinguish one network from another. Accuracy, uniqueness and current records support coordination.

The record does not operate BGP. It does not store the complete route table of a Comcast router, select a path, enforce a table threshold or direct Chicago traffic away from California. Those outcomes arise from running software, installed state, topology and operator policy.

This distinction avoids two errors. The first is treating registry attribution as proof of every internal action. Seeing AS7922 in a path can identify a network context, but it cannot identify the employee, configuration or legal responsibility behind a failure. The second is dismissing registry evidence because it cannot enforce forwarding. A current record remains useful for attribution and incident coordination even though it is not a path-control system.

Network accountability depends on connecting the record layer to the reality layer. The resource identifier, device inventory, routing policy, table telemetry, change record, alarm and external path observation should refer to the same operational event. When those links are preserved, a review can ask who controlled each decision without pretending that one database governed the entire network.

Reporting accountability and confidential evidence

The FCC's Part 4 rules and related guidance establish reporting and recordkeeping requirements for qualifying communications outages.[7][8] The rules recognize that geographic scope, duration, users and public-safety effects can matter to oversight. They also protect outage information that is not generally public.

This article does not have Comcast's confidential NORS filing for the November 2021 events. It therefore cannot state what Comcast reported as the root cause, how many users were counted under regulatory definitions, whether a threshold was met, or what remediation was supplied to the FCC.

That boundary matters because a filing and a public postmortem serve different audiences. A regulator may receive sensitive infrastructure detail that should not be publicly exposed. Customers and dependent operators still need enough public information to understand the nature of a failure and evaluate continuity. Accountability does not require publishing exploit-ready topology. It does require a credible public explanation of the failure class, scope, recovery and measurable prevention.

A useful public record could state that a routing-state limit was exceeded, describe the affected control domain at an appropriate level, give the observation and recovery windows, explain why rerouted traffic entered the same domain, and list the controls changed afterward. It could do so without naming individual engineers or publishing sensitive router configurations.

The absence of such details in the frozen public packet limits conclusions. It does not prove that Comcast failed to report confidentially or failed to remediate. It shows that outsiders must rely primarily on external measurements for the technical reconstruction.

Control owners and evidence obligations

Accountability should follow practical control rather than proximity to a headline.

Comcast controlled internal routing capacity. The operator could inventory platforms, set or accept limits, monitor occupancy, reserve headroom and test failure behavior. Evidence would include device and software inventories, table telemetry, limit configurations, alarms and capacity-test results.

Comcast controlled topology and route distribution. It could design core failure domains, route-reflector relationships, path preferences and geographic constraints. Evidence would include approved topology, routing policy, dependency maps and fault-injection results. The public packet does not disclose those materials.

Comcast controlled change and recovery. It could authorize changes, stage them, preserve before-and-after state, execute rollback and verify forwarding. Evidence would include tickets, approvals, diffs, command logs, incident decisions and external recovery tests.

Suppliers controlled product behavior within their products. A router or software supplier may define capacity, alarms and failure modes. The frozen sources do not identify a supplier or product, so the article cannot assign a supplier-specific obligation or defect.

External observers controlled measurement quality. ThousandEyes controlled its vantage points, tests, path interpretation and published analysis.[1]-[4] Its evidence can show patterns but should disclose limits and remain open to comparison with internal data.

Customers controlled only their own continuity choices. An enterprise may use multiple access providers, paths or application regions. Those options can reduce dependency, but they do not transfer responsibility for Comcast's internal routing state to the customer. Some residential or public-service users may have no practical substitute.

ARIN controlled registry accuracy and availability for its records. It did not control Comcast's internal routes.[9]

The FCC controlled reporting rules and protected oversight records. It did not control the routing decision that sent a path through Sunnyvale.[7][8]

This distribution is not a claim that every actor failed. It is a map of who could produce the evidence needed to evaluate a particular control.

Measurable remediation

The strongest remediation program converts the incident's unknowns into recurring tests.

1. Define every relevant limit. For each routing and forwarding structure, record the platform maximum, configured maximum, current occupancy, expected growth and emergency reserve. Separate received, accepted, selected and installed state where the platform exposes them. A single headline route count is insufficient if a smaller internal structure can fail first.

2. Set multi-stage alarms. Warning thresholds should leave enough time for investigation before the hard limit. Alerts should include the affected structure, current value, rate of change, relevant neighbor or process and safe response. Alarm delivery must be tested over an independent management path.

3. Test transient headroom. Capacity models should include normal growth plus the extra state created during maintenance, route reconvergence, session restoration and policy rollback. The test should measure the peak, not only the final stable table.

4. Verify failure behavior. In a controlled environment, approach or exceed the configured limit and record what the system does. Does it reject new routes, reset a session, withdraw existing routes, restart a process, preserve forwarding or produce churn? A documented limit without a verified failure mode is incomplete.

5. Map shared control dependencies. Redundant nodes should be checked for common route reflectors, configuration systems, software versions, table ceilings, power domains and management paths. A failover target that shares the same limiting resource is not independent.

6. Test geographic locality. Define which classes of traffic should remain inside a region during specified failures. Use internal and external measurements to verify that a local path is not redirected through a distant impaired core without an explicit and tested reason.

7. Couple control-plane and forwarding evidence. A route can exist in a control table while packets still fail. Recovery should require successful forwarding tests, acceptable loss and stable paths from multiple vantage points. Internal route state and external path evidence should be time-aligned.

8. Preserve trigger, detection, response and recovery separately. The initiating event may precede the first alarm. An external observer may detect symptoms before the operator identifies the cause. The rollback may occur before global paths stabilize. A postmortem should record each timestamp and evidence source rather than compressing them into one outage duration.

9. Review route-reflector and convergence behavior. Where route reflection is used, test client behavior, alternate reflector capacity, path visibility and state repopulation.[12] Define acceptable convergence and loss for each planned fault.[11] Graceful mechanisms can be evaluated where relevant without assuming that they solve table exhaustion.[13][15][19]

10. Keep interdomain policy terminology precise. RFC 7908 defines route leaks, while RFC 8212 and RFC 9234 address explicit policy and relationship-aware controls.[17][18][20] The Comcast evidence in this packet concerns internal path changes and a table limit. Operators should not label every unexpected reroute a route leak, because an incorrect label points remediation at the wrong control.

11. Exercise the incident-management dependency chain. The monitoring system, authentication service, status page, customer support, engineering communication and change path should remain available when the production core is impaired. The exercise should include a loss of primary network reachability.

12. Publish a bounded technical account. A public report should identify the failure class, time window, affected control domain, recovery method and verified remediation without disclosing sensitive topology. It should distinguish measured facts, internal findings and unresolved questions.

Each measure needs a pass condition. "Monitor table size" is not a pass condition. "Alert at a defined reserve, deliver the alarm over an independent path within a tested interval, and demonstrate that operators can restore adequate headroom before the hard limit" is testable. "Provide redundancy" is not a pass condition. "Under an isolated Sunnyvale-core failure, specified traffic remains reachable without traversing the isolated domain and with loss below a defined threshold" is testable.

The controls should also have owners and review dates. Capacity may change as customers, peers, prefixes, services and traffic-engineering policies change. A passing result from one year cannot establish indefinite safety. The evidence should show when the test was run, against which software and topology, and which exceptions remain.

These proposals are not proof that Comcast omitted them. They are the measurable controls logically connected to the externally observed failure and the attributed limit mechanism.

Later standards are context, not verdicts

Several IETF documents in the source set postdate or generalize beyond the incident. RFC 8212 describes default-reject behavior when external BGP policy is not explicitly configured.[18] RFC 9234 describes BGP Roles and the Only-to-Customer attribute for reducing certain route leaks.[20] Neither document establishes the cause of an internal routing-table limit or proves Comcast's 2021 configuration.

RFC 7454 collects BGP operational and security practices.[16] RFC 6198 and RFC 8326 address graceful shutdown requirements and signaling.[15][19] RFC 5880 defines BFD.[14] They help frame questions about policy, planned change and detection. They should not be presented as a checklist that the public evidence proves Comcast violated.

The distinction protects the analysis from hindsight bias. Standards can show that a control was known or technically possible. They cannot establish that it was contractually required, supported on a specific platform, configured in the affected network, or capable of preventing the exact failure. Those conclusions require additional evidence.

The article uses standards to define measurable alternatives. It does not use them to manufacture a finding of fault.

What this article does not conflate

The November 2021 events are not Comcast's 2017 outage associated with a Level 3 BGP route leak. That event involved leaked external routes and a different observed mechanism. It is not Comcast's 2018 fiber-cut outage, where physical damage and more-specific announcements formed a different recovery record. It is not the 2023 CitrixBleed-linked customer-data incident, which concerned an exposed edge appliance and identity records.

The article also does not equate all rerouting with a route leak. A route leak has a specific interdomain policy meaning.[17] The 2021 packet shows internal or provider-controlled paths changing and traffic entering an impaired core. Without the relevant route advertisements and relationship evidence, calling that behavior a route leak would be unsupported.

The outage is not evidence that a registry failed. ARIN's AS7922 record helps identify the network resource.[9] A correct record cannot prevent a route table from reaching a limit, and a table limit does not make the registry record incorrect.

Finally, the outage is not evidence of an attack. The frozen sources do not establish malicious action, unauthorized access, deliberate disruption or criminal intent.

Essential uncertainties

The exact table and threshold remain unknown. The triggering state change, device, software, command and responsible role remain unknown. The topology and table occupancy of each Sunnyvale node remain unknown. The complete customer and service impact remains unknown.

The internal alarm chronology, response decisions, recovery method and post-incident remediation are not in the frozen public packet. The contents of any confidential outage filing are unavailable. No source used here establishes a supplier defect, an individual failure, negligence, legal liability or quantified loss.

These gaps limit blame claims, not operational questions. The observed paths and attributed limit still identify which evidence would be necessary for a complete review.

Conclusion

Comcast's November 2021 outages turned a capacity boundary into a network-accountability test. ThousandEyes observed two incidents centered on the Sunnyvale core. Traffic already using the core failed, and some traffic that had been successful failed after rerouting into Sunnyvale. During the second event, paths from distant regions were directed through the same area and sometimes alternated between loss and reachability.[1][3]

A later ThousandEyes review attributed the incident to an inadvertently exceeded routing-table limit.[2] That explanation does not disclose the exact internal trigger, but it defines a concrete control surface. Routing-table occupancy, thresholds, transient headroom, failure behavior, topology dependencies, alarms, rollback and forwarding-plane recovery can all be measured.

The event also shows why records and diagrams are not enough. ARIN can accurately record AS7922.[9] BGP can calculate a new path.[10] A topology can contain multiple nodes. Continuity still fails if running state reaches a limit or if an alternate path enters the same impaired domain.

Accountability therefore rests on evidence of operational control. Comcast controlled the routing system and its recovery. External observers controlled their measurements. Regulators controlled reporting requirements. Suppliers may have controlled product behavior, but the public packet does not identify one. Customers controlled only the alternatives practically available to them.

The defensible conclusion is neither that one standard would have prevented the outage nor that an operator can never fail. It is that a continuity claim should be backed by tested headroom, independent failure domains, durable monitoring, bounded public explanation and externally verified recovery. In a large network, resilience is not the presence of another route. It is proof that the other route works when the primary control domain does not.

Sources

  1. https://www.thousandeyes.com/blog/comcast-outage-analysis-nov-9-2021
  2. https://www.thousandeyes.com/blog/seven-outages-shook-up-2021
  3. https://www.thousandeyes.com/blog/internet-report-weekly-pulse-nov-15
  4. https://www.thousandeyes.com/blog/analyzing-internet-issues-traffic-outage-detection
  5. https://corporate.comcast.com/comcast-voices/one-of-the-most-sophisticated-networks-in-the-world-2
  6. https://corporate.comcast.com/press/releases/comcast-harnessing-cloud-and-ai-to-transform-next-generation-internet-experiences
  7. https://docs.fcc.gov/public/attachments/DA-22-1300A1.pdf
  8. https://www.ecfr.gov/current/title-47/chapter-I/subchapter-A/part-4
  9. https://rdap.arin.net/registry/autnum/7922
  10. https://www.rfc-editor.org/rfc/rfc4271
  11. https://www.rfc-editor.org/rfc/rfc4277
  12. https://www.rfc-editor.org/rfc/rfc4456
  13. https://www.rfc-editor.org/rfc/rfc4724
  14. https://www.rfc-editor.org/rfc/rfc5880
  15. https://www.rfc-editor.org/rfc/rfc6198
  16. https://www.rfc-editor.org/rfc/rfc7454
  17. https://www.rfc-editor.org/rfc/rfc7908
  18. https://www.rfc-editor.org/rfc/rfc8212
  19. https://www.rfc-editor.org/rfc/rfc8326
  20. https://www.rfc-editor.org/rfc/rfc9234