Summary
- Frozen event boundary: On 6 November 2017, independent monitoring linked a major Comcast reachability disruption to routing changes associated with Level 3 Communications and AS3356. ThousandEyes placed its principal observed Comcast impact roughly between 09:45 and 11:25 Pacific, with routing inconsistencies visible from about 09:30. [2] Contemporary reports carried Level 3's statement that a configuration issue caused disruption. [1][5][6] This article does not merge the event with Level 3's separate 2016 outage, CenturyLink's December 2018 911 outage or the August 2020 AS3356 FlowSpec incident.
- Observed mechanism: ThousandEyes reported more-specific announcements for over one thousand Comcast subsidiary and customer prefixes and observed traffic that previously crossed Comcast's AS7922 backbone being directed through Level 3 AS3356, with increased packet loss and latency. [2] That is strong evidence for a routing-policy failure and its forwarding consequences. It does not reveal the exact internal command, route-policy entity, deployment process or complete set of affected users.
- Responsibility boundary: Level 3 controlled its configuration generation, approval, deployment scope, BGP export policy, route monitoring, rollback authority and incident communication. Peers and downstream networks controlled import policy, customer-cone expectations, prefix limits, anomaly detection and failover choices. Comcast controlled subscriber communication and parts of the user-facing recovery path. Customers and end users could observe failure but could not repair the interdomain routing state.
- Control lesson: A configuration can be syntactically valid and still violate an operational invariant. Backbone change control must test intended reachability, export scope, path changes and blast radius before deployment, then compare live route state and user reachability against those expectations. Batfish used the event to argue for model-based network validation rather than confidence in manual review alone. [3]
- Reality layer: Registry and ASN records help identify resources and responsible operators, but they do not enforce BGP policy. Route collectors show what running networks announced and accepted; forwarding probes show whether traffic reached the intended destination. Accountability depends on connecting intended policy, running announcements, packet paths, rollback records and user-visible restoration.
The incident has to remain bounded to 6 November 2017
The first requirement of an accountable reconstruction is to define which event is under examination. Level 3 and the later CenturyLink or Lumen organization appear in several major outage records. Combining those records would produce a longer narrative but a weaker finding because the mechanisms, control owners and evidence differ.
This article concerns the BGP event observed on Monday, 6 November 2017. ThousandEyes reported that its employees using Comcast connections encountered failures involving services such as Slack, Gmail and Webex. Its measurements placed the major Comcast impact from approximately 09:45 to 11:25 Pacific. It also identified BGP inconsistencies from around 09:30, before the full user-visible interval it described. [2]
Wired reported that a Level 3 configuration error affected portions of US internet access and that the provider said service was restored after the issue was corrected. [1] Other contemporaneous reports described widespread complaints across several access providers and carried a Level 3 explanation centered on a configuration issue. [5][6] Those reports establish a broad public impact signal, but complaint maps and simultaneous reports do not prove that every network failed for the same technical reason.
ThousandEyes provides the strongest public bridge between route state and user experience. Its analysis described packet loss inside Comcast and Level 3 paths, route changes for Comcast-related destinations and a return toward normal after Level 3 withdrew the leaked routes. [2] APNIC's later annual routing-security review also identified the incident as one of 2017's notable large-scale BGP events. [4]
Three exclusions are essential.
First, the incident is not Level 3's 2016 outage. That separate event later drew regulatory and industry scrutiny and involved a different mechanism. [8] Second, it is not the December 2018 CenturyLink outage that disrupted transport and 911 service. Third, it is not the August 2020 AS3356 FlowSpec incident, in which a traffic-filtering action propagated through a backbone and caused a different failure pattern.
The shared operator history can support comparative governance questions. It cannot substitute one event's evidence for another. The 2017 conclusion must stand on the 2017 route observations, provider statement and contemporaneous user-path evidence.
AS3356 made a configuration change a cross-network event
An autonomous system is a network or group of networks presenting a common routing policy to the internet. BGP lets autonomous systems exchange reachability information, including the prefixes they can reach and the sequence of autonomous systems represented in a path. RFC 4271 defines the protocol mechanics and the decision framework within which operators apply local policy. [14]
Level 3 operated AS3356, a major global transit network. RIPEstat records provide a public resource and routing view for that ASN. [10] An ASN record helps identify the operator associated with route observations. It does not show every private peering contract, customer relationship or router policy active during an incident.
Scale changes the accountability of configuration. A small enterprise may advertise a limited set of prefixes to one provider. A large transit network exchanges routes with access providers, content networks, enterprises and other carriers. A policy error at that position can influence paths far beyond the operator's direct retail customers.
The internet does not have one central controller that approves every route. Each network decides what to announce and what to accept. That design supports independent operation, but it also means a bad announcement can propagate when multiple networks consider it acceptable under their local policies. The consequences depend on prefix specificity, path attributes, business relationships, filters and the timing of route selection.
ThousandEyes described Level 3 announcing more-specific routes associated with Comcast subsidiary networks and customers. Traffic that had gone through Comcast's AS7922 backbone was observed going through AS3356 instead. [2] More-specific routing matters because forwarding generally follows the longest matching prefix. An announcement covering a narrower address block can attract traffic even when a broader route remains visible.
The evidence does not mean that one router dictated the whole internet. It means a network with broad interconnection originated or propagated information that enough other networks accepted to produce substantial forwarding changes. Each acceptance was a local policy decision, but Level 3 controlled the change that introduced the observed route state.
This is why a backbone configuration cannot be treated as an internal administrative detail. Its practical output is a set of assertions consumed by other autonomous systems. The larger the interconnection footprint, the stronger the obligation to model export scope, stage deployment and verify the external effect.
BGP evidence and forwarding evidence answer different questions
Route collectors and user-path probes observe related but different parts of the incident.
A route collector records BGP updates from participating peers. RouteViews archives provide historical update data for November 2017. [11] RIPE NCC's Routing Information Service similarly collects routing information from vantage points around the world. [12] CAIDA BGPStream provides tools and data interfaces for analyzing BGP events. [13]
These systems can help reconstruct when a prefix was announced or withdrawn, which origin and path appeared at a collector, and how visibility changed over time. They cannot see every route in every router. A route that was not visible at one collector may still have existed elsewhere. A visible route may not have carried substantial traffic. Local preference, traffic engineering and private peering can produce forwarding behavior not fully represented in public collector feeds.
Forwarding and application measurements add another layer. ThousandEyes reported path changes, latency and packet loss from distributed agents. [2] Those measurements show what selected packets experienced from selected locations. They are closer to user impact than a control-plane update alone, but they still do not represent every user or path.
An accountable investigation joins the layers:
- the intended route policy before the change;
- the exact configuration diff or generated policy;
- BGP announcements and withdrawals seen internally;
- announcements observed by independent collectors;
- forwarding paths from diverse networks;
- packet loss, latency and successful transactions;
- customer complaints and provider status updates;
- the rollback or corrective change;
- independent confirmation of restored routing and reachability.
Without this join, teams can misclassify the event. A route update may be visible without causing material harm. Packet loss may occur without a BGP cause. An application can fail because DNS, identity or a cloud dependency failed while its route remained stable.
The 2017 evidence is persuasive because route and forwarding observations point in the same direction. ThousandEyes saw route inconsistencies and a changed AS path together with increased latency and packet loss. [2] The provider attributed the disruption to configuration. [1][5][6] That combination supports a BGP misconfiguration conclusion while preserving uncertainty about internal implementation.
The same discipline applies to recovery. A withdrawal seen at one collector is not enough to declare every user path restored. Operators should confirm stable route state, expected path selection, reduced loss, normalized latency and successful application transactions across relevant regions.
Route leak is more precise than hijack, but still requires attribution
Public discussion often uses route leak and route hijack interchangeably. The distinction matters because it shapes claims about intent, authorization and prevention.
RFC 7908 defines a route leak as the propagation of routing announcements beyond their intended scope. [16] The document describes categories based on relationships among networks and the direction in which routes are propagated. A leak can occur when a customer exports provider-learned routes to another provider, when internal routes escape or when a network announces information contrary to an intended business relationship.
A hijack often refers to an unauthorized origin or path that attracts traffic, sometimes maliciously. The public evidence for 6 November 2017 supports an inadvertent configuration failure, not an assertion of malicious intent. Contemporary reporting described a configuration issue, and ThousandEyes used route-leak terminology. [1][2][5][6]
Even the leak label should be tied to evidence. Public observers did not have Level 3's full policy intent, contracts and router configurations. ThousandEyes observed more-specific announcements and path changes inconsistent with normal Comcast routing. [2] That behavior is consistent with a leak, and named analysts characterized it that way. The article should preserve that attribution rather than claim access to private intent.
Terminology also affects control analysis. If the central problem is unauthorized origin, route-origin validation may address part of it. If the origin remains authorized but export scope violates a relationship policy, origin validation can still report the route as valid. If a more-specific route is technically covered by authorization but operationally wrong, acceptance requires other policy controls.
The safest conclusion is therefore bounded: Level 3 acknowledged a configuration issue; independent measurements observed AS3356-related route changes and forwarding impairment; analysts characterized the event as a route leak. The sources do not establish sabotage, credential compromise, interception intent or one exact internal command.
That boundary is not evasive. It prevents technical uncertainty from becoming an accusation. It also keeps remediation focused on the controls the evidence supports: change validation, export policy, peer filtering, route monitoring and rollback.
A syntactically valid configuration can still be operationally wrong
Network change systems often check whether configuration text parses and whether a device accepts it. Those checks are necessary, but they do not prove that the resulting network behavior matches policy.
A BGP configuration can be valid to a router while violating an operational invariant. It may export a route to the wrong neighbor, accept a customer route outside an expected set, create a more-specific path with unintended reach, alter preference or remove a filter. The router executes the command it received. The failure lies in the gap between syntactic validity and intended network state.
Batfish's analysis of the Level 3 event argued for testing network configuration against intended properties before deployment. [3] Model-based validation can ask whether important destinations remain reachable, whether forbidden paths appear, whether routes escape their intended scope, whether redundancy survives a failure and whether a change affects more devices or prefixes than expected.
This does not mean a model can perfectly reproduce the public internet. External peers have private policies, route state changes continuously and some relationships are not documented. A useful model is explicit about those limits.
The change-control chain should therefore contain several checks:
- Source control: The proposed configuration or policy entity is stored as a reviewable diff with an identified owner.
- Schema and syntax validation: Tools confirm the configuration is accepted and references valid entities.
- Policy invariants: Automated tests check export scope, accepted prefixes, expected origins, path constraints and reachability.
- Blast-radius calculation: The system estimates affected routers, sessions, prefixes and customer classes.
- Representative staging: The change is tested against topology and policy state close enough to production to expose meaningful conflicts.
- Canary deployment: A limited, observable subset receives the change before wider rollout.
- Independent telemetry: Route collectors, peer views and forwarding probes compare intended and observed effects.
- Automatic stop conditions: Unexpected route count, path change, loss or latency blocks further deployment.
- Rollback authority: A named operator can reverse the change without waiting for a long approval chain.
- Post-change verification: The team proves expected routes and services remain stable.
No single check eliminates risk. Together they reduce the chance that one administrative error becomes a backbone-wide incident.
Accountability requires evidence that these controls existed and operated. A postmortem saying a configuration error occurred does not answer whether the change had peer review, whether tests covered export behavior, what alarms fired or how quickly rollback was authorized.
Blast radius should be a pre-deployment property
Operational teams often describe blast radius after an incident by counting affected services, prefixes or users. For high-consequence network changes, blast radius should also be estimated before deployment.
The question is not only how many devices receive a configuration. A change applied to one route-policy entity can affect many BGP sessions. A change on one edge router can alter advertisements consumed by a large peer. A more-specific prefix can redirect traffic without a large device count. Shared templates can turn one line into a fleet-wide behavior.
A pre-deployment blast-radius assessment should ask:
- which routers and sessions reference the changed entity;
- which prefixes can match the policy;
- which neighbors can receive new or changed announcements;
- whether the change affects customer, peer and provider relationships differently;
- which critical services depend on the affected routes;
- whether rollback itself will generate a large update burst;
- whether monitoring covers the likely external paths;
- whether one canary provides a meaningful sample of the final deployment.
The assessment should include uncertainty. If peer policies are unknown, that uncertainty is a reason to narrow the initial deployment and strengthen external monitoring. It is not a reason to assume peers will contain the error.
ThousandEyes reported over one thousand more-specific routes in the 2017 event. [2] APNIC's annual review placed the incident among large routing-security events affecting thousands of autonomous systems. [4] These figures are observations from named analyses, not a complete internal count. They nonetheless show why route quantity and external propagation should have been stop conditions.
The operational objective is not to guarantee zero route updates. Networks must change. The objective is to make the intended scope measurable and to detect when the running state departs from it.
For a backbone, the stop condition may combine route count, new origin or path patterns, peer-specific export changes, collector visibility, packet loss and customer alarms. A threshold should be linked to the change request so responders know whether an anomaly is expected, tolerable or grounds for rollback.
When an incident occurs, the same blast-radius model becomes part of the evidence record. Investigators can compare predicted and actual scope, identify missing dependencies and improve the next test.
Rollback is a production capability, not a sentence in a plan
A change-control process is incomplete if rollback exists only as an instruction to restore the previous configuration. BGP rollback can itself produce convergence, withdrawal and traffic shifts. It has to be designed and tested as an operational action.
The public record says Level 3 corrected the configuration issue and ThousandEyes observed withdrawal of the leaked routes by about 11:25 Pacific. [1][2] It does not disclose who authorized the rollback, whether the previous configuration was restored atomically, how devices converged or which external signals were used to confirm recovery.
Those unknowns define the evidence an accountable operator should retain:
- the change identifier and exact diff;
- deployment start, scope and operator;
- first anomaly and alert;
- incident declaration and command owner;
- decision to stop or reverse rollout;
- rollback command or replacement policy;
- completion by device and session;
- route withdrawal and expected reannouncement;
- forwarding recovery by region and peer;
- customer and access-provider confirmation;
- post-rollback stability interval.
Rollback speed is not the only measure. A rapid reversal that leaves stale route state or overloads sessions may prolong harm. A slower staged rollback may be justified if it prevents a second failure. The record should explain the decision and show its effect.
The team also needs an out-of-band path to control the network if production routing is degraded. Management access that depends on the same path under repair can convert a routing error into a recovery failure. The 2017 public evidence does not say that Level 3 lost management access. The point is a control requirement derived from the failure class, not an incident-specific assertion.
Customers need a parallel rollback plan. An enterprise that sees an upstream route failure may shift traffic, change announcements or invoke another provider. Such actions can create their own propagation risks. The customer should define who can act, which paths are independent and how to verify that a failover does not worsen the event.
Rollback becomes credible when operators can demonstrate it in exercises and show that live incident evidence matches the planned stages. A generic statement that configuration was corrected is a starting fact, not complete proof of recovery governance.
Peers had containment controls of their own
The network that introduced a bad route is the primary control owner for the change, but interdomain routing distributes responsibility. Each peer decides what it accepts, prefers and propagates.
RFC 7454 describes operational security practices for BGP, including prefix filtering, AS-path filtering, limits and relationship-aware policy. [15] MANRS similarly sets expectations for filtering, coordination, global validation and anti-spoofing among network operators. [18]
Peer responsibility is not equal in every relationship. A provider has stronger grounds to know the expected prefix set and routing role of a customer than of a settlement-free peer. A customer may not have a complete list of every provider route. Large dynamic networks make static filters harder to maintain. Policy errors can occur on either side.
Still, a network accepting routes should be able to explain its trust model:
- which prefixes and origins are expected from each customer;
- whether more-specific announcements are permitted;
- whether an AS path is consistent with the relationship;
- what route-count or max-prefix limits apply;
- whether unusual announcements trigger alarms or rejection;
- how exceptions are approved and expired;
- which independent data sources validate expectations;
- how emergency coordination occurs with the announcing network.
Public registry information can support these controls, but it may be stale or incomplete. Internet Routing Registry entities, RPKI authorizations and observed paths each answer different questions. Operators should not treat one data source as a complete policy oracle.
The 2017 event shows a collective containment problem. Level 3 controlled the configuration that produced the observed announcements. Other networks accepted enough of those announcements to change traffic paths. Some may have had valid operational reasons under the information available to them. Others may have lacked filters that could have limited propagation.
An accountable review should not assign blame to unnamed peers without their configurations. It should ask which containment controls were technically available, which were in use, what alerts fired and whether later exercises prove improvement.
The same logic protects against cost transfer. A backbone can externalize the effect of a configuration error to access providers, content networks and users. Peers can externalize weak filtering to the wider routing system. Shared evidence and coordinated remediation are necessary because no single operator controls every acceptance decision.
RPKI helps with origin authorization but not every leak
RPKI lets address holders create Route Origin Authorizations that identify which autonomous systems may originate specified prefixes. Route Origin Validation classifies a route according to whether its origin and prefix length are consistent with a valid authorization. RFC 6811 defines that validation state and how it can inform local policy. [17]
This is an important control, but its scope must be stated precisely.
If an unauthorized autonomous system originates a prefix, a valid ROA can help networks identify and reject the invalid route. If a route leak preserves an authorized origin but violates export scope, the route can remain origin-valid. RPKI origin validation does not encode the complete business relationship or intended path.
The Level 3 public record does not establish the ROA state for every affected prefix in November 2017, the ROV policy of every peer or a counterfactual in which one control would have prevented the incident. The article therefore does not claim that RPKI would have stopped the event.
RPKI belongs in the remediation framework as one layer:
- ROAs make origin authorization explicit;
- ROV can reject some unauthorized origins;
- prefix filters limit what a neighbor may announce;
- relationship-aware path policy limits route propagation;
- max-prefix controls limit quantity;
- anomaly detection identifies unexpected changes;
- model-based change validation tests intended export behavior;
- route collectors and forwarding probes verify live outcomes.
Later mechanisms such as BGP Roles and route-leak prevention can encode parts of the relationship boundary more directly. They should be evaluated as later controls, not retroactively asserted as deployed in 2017.
The broader accountability point is that a security control must match the failure mode. Labeling all BGP problems as an RPKI gap can create false assurance. A network can have complete origin authorization and still export valid routes in the wrong direction or with an unintended more-specific.
Operators should report the control they believe would have interrupted the actual chain. If the failure was policy generation, show a new invariant test. If the failure was a missing peer filter, show the filter and route-rejection exercise. If an invalid origin propagated, show ROV coverage. Each claim should be linked to observed behavior.
Registry records are evidence, not route enforcement
The Heng.lu doctrine draws a useful distinction between a record and a running system.
ASN and address records preserve identifiers, resource holders, contacts and registration history. Routing registries can record intended policy. RPKI can record origin authorization. These systems support uniqueness, traceability, transfer records, security metadata and coordination.
They do not forward packets or enforce every peer's policy by declaration.
RIPEstat can help an investigator identify AS3356 and view public routing data. [10] RouteViews, RIPE RIS and BGPStream can show announcements observed by collectors. [11][12][13] Those records are part of the accountability ledger. The running route accepted by each network and the forwarding path selected for packets remain the reality layer.
This distinction prevents two errors.
The first error is treating registration as proof of healthy operation. A correctly registered ASN can announce routes under a faulty policy. Accurate resource data does not prove that a configuration change is safe.
The second error is treating registry administrators as sovereign controllers of interdomain routing. Operators choose local policy and run the routers. Recordkeepers can improve evidence and security metadata, but they do not replace operational responsibility.
For the 2017 incident, the evidence chain should therefore connect:
- the registered ASN and address resources;
- the intended Level 3 and Comcast relationship policies;
- the configuration diff that changed route behavior;
- internal route state;
- collector-observed announcements;
- peer acceptance;
- actual forwarding paths;
- user-visible loss and latency;
- withdrawal and restoration.
No one layer is enough. A private configuration diff without external evidence may miss propagation. A collector update without intent cannot prove why the route appeared. A user complaint without route data cannot identify the failing control.
The doctrine is a practical governance rule, not advocacy copy. It says that accountability should follow the parties operating the running system while accurate records preserve who controlled resources and what policy was claimed.
Incident communication should name the failing layer
Users experiencing the 2017 disruption saw applications fail or slow down. They generally did not see a BGP policy entity, an AS path or the route that redirected traffic.
That gap makes incident communication part of the technical control system. A statement that there is an "internet outage" is too broad to guide an access provider, enterprise or content network. A statement that a configuration issue affected routing narrows the problem. A useful update goes further without exposing sensitive details.
An accountable notice can state:
- the affected network layer;
- the approximate start time and detection source;
- the broad scope known at that time;
- whether the operator stopped further changes;
- whether routes are being withdrawn or restored;
- which customer classes or peer regions remain affected;
- what evidence will define restoration;
- which facts remain unconfirmed.
The public reports carried Level 3's configuration-issue explanation. [1][5][6] That was more informative than an unexplained service-degradation notice. The public record does not show a detailed provider postmortem with the full internal sequence and control changes.
Access providers also had a communication duty. Comcast users saw service failures, and ThousandEyes' analysis centered on Comcast paths. [2] Comcast controlled the customer relationship and could describe subscriber impact even though it did not control Level 3's configuration.
Communication should preserve uncertainty. Reports of AT&T, Verizon, Spectrum and other networks experiencing problems at similar times do not prove a common technical cause. [5][6] A provider should separate confirmed shared impact from correlated complaints still under investigation.
The restoration message also needs evidence. "Resolved" should mean more than the change was rolled back. It should be supported by stable routes, expected paths, normalized packet loss and successful customer transactions over a defined interval.
Precise communication reduces operational cost. It helps customers decide whether to fail over, preserve logs or wait for upstream recovery. It also creates a timestamped record against which later claims can be tested.
User impact should not be inflated beyond measured paths
Contemporary coverage described a widespread or nationwide disruption. ThousandEyes saw effects across several US regions and reported that millions of Comcast users were potentially involved. [1][2][5][6] These accounts establish material impact. They do not justify a claim that every Comcast customer or every reported provider suffered the same outage.
Impact assessment should distinguish:
- route visibility at collectors;
- forwarding changes from specific probes;
- packet loss and latency;
- inability to reach named destinations;
- access-provider service complaints;
- application transaction failures;
- duration by geography and network;
- potential users versus confirmed failed sessions.
One person may have had a cached DNS answer and a working path while another could not reach the same service. One enterprise may have used a second transit provider. A mobile application may have retried through a different endpoint. The same BGP event can produce heterogeneous outcomes.
The source packet does not contain a defensible aggregate financial-loss figure. It also does not establish legal causation for every business interruption. This article does not manufacture one by multiplying a user estimate by an outage duration.
An operator with internal telemetry can do better. It can report traffic shifts, dropped packets, failed sessions, affected prefixes, customer tickets and restoration by region. Access providers can measure subscriber-session and application-level effects. Large customers can measure failed transactions and dependency-specific losses.
These measures should be reconciled rather than collapsed into one headline number. Route count measures network state. Packet loss measures a path symptom. Complaint volume measures visible frustration. Transaction failure measures business effect. Each is useful when its denominator and limits are stated.
The restraint matters for accountability. Inflated claims can make the report easier to dismiss. Bounded measurements identify where controls failed and what remediation must prove.
Independent reconstruction has limits that should be documented
Public route data is unusually valuable because it allows an incident to be studied outside the operator. That independence creates accountability, but it does not create omniscience.
RouteViews and RIPE RIS see the routes sent to their collectors by participating peers. [11][12] BGPStream helps researchers process and compare those observations. [13] RIPEstat combines resource and routing views. [10] ThousandEyes adds forwarding and service tests from distributed vantage points. [2]
Together, these sources can establish:
- that selected routes changed;
- which origins and paths were visible;
- approximate timing of announcements and withdrawals;
- whether traffic followed a changed path from measured locations;
- whether loss or latency rose;
- when measured reachability recovered.
They generally cannot establish:
- the exact command entered by an operator;
- the internal policy-generation chain;
- every route in every router;
- private peering and local-preference decisions;
- the complete customer set;
- the decision owner for rollout or rollback;
- internal alarms and incident communications;
- current effectiveness of remediation.
An independent report should label collector coverage, clock precision, normalization choices and missing data. It should preserve raw update references where possible so another analyst can reproduce the finding.
The operator should retain a richer package. That package can include configuration versions, device logs, route-policy evaluations, route-reflector state, peer notifications, telemetry, packet-path tests, incident tickets and change approvals. Sensitive details can be shared with auditors or affected partners under controls.
The most credible postmortem connects the private and public views. It explains why the observed announcements appeared, which internal control failed, how the route was contained and what test now prevents recurrence.
Where the operator does not publish that evidence, independent observations still support a bounded conclusion. They should not be stretched to fill private gaps.
Remediation must be testable against a repeatable route fault
The public record reviewed here does not establish every remediation Level 3 or its peers implemented after November 2017. A responsible assessment therefore defines what evidence would demonstrate repair rather than asserting current failure or success.
A testable remediation program should cover five control groups.
Configuration generation and review
The operator should show that routing policy is generated from controlled data, reviewed as a diff and checked against invariants. The evidence should identify which route classes and relationships a change can affect.
Staged deployment and stopping conditions
The operator should deploy through a bounded canary where possible, compare live announcements with intent and stop when route count, path changes, loss or latency exceeds the approved boundary.
Peer and customer filtering
The operator and peers should maintain expected prefix and relationship policies, test exception handling and monitor for unexpected more-specific announcements. Controls should use multiple evidence sources rather than assuming one registry is complete.
Rollback and recovery
Teams should exercise rollback, route withdrawal, session convergence and out-of-band management. Recovery should be confirmed by independent route and forwarding observations.
Disclosure and verification
Post-event reporting should distinguish confirmed cause, contributing controls, scope, unknowns and remediation. Later tests should show whether the new control detects or contains a representative fault.
A realistic exercise could introduce a safe synthetic policy error in a lab or isolated route domain. The error should attempt to export a more-specific route beyond an intended relationship. The system should reject it at generation or pre-deployment validation. If it reaches a canary, monitoring should stop rollout. A peer test should demonstrate import filtering. A rollback test should remove the route and confirm expected forwarding.
The exercise should not inject harmful routes into the public internet. Its purpose is to prove the control chain in a representative environment and preserve auditable evidence.
Remediation is strongest when tied to the original failure path. A generic investment in monitoring does not prove export-policy validation. A new RPKI program does not prove containment of origin-valid leaks. A revised procedure does not prove that rollback works under route churn.
What boards and service buyers should ask
Backbone routing risk can appear too technical for board oversight. The relevant governance questions are concrete.
Boards should ask how many changes can alter externally advertised routes, who can approve them, what pre-deployment invariants are mandatory and what automatic stop conditions exist. They should request evidence from exercises, not a maturity label.
Service buyers should ask whether their providers monitor routes and forwarding from outside their own network, how quickly they notify customers of routing incidents and whether alternate paths are operationally independent. A contract naming two carriers does not prove route diversity if both depend on one backbone or shared facility.
Network operators should ask whether registry, RPKI and observed-route data are reconciled, how customer-cone expectations are maintained and how exceptions expire. They should know which peers can send unusually broad or more-specific route sets.
Incident leaders should ask who can freeze deployment, who can authorize rollback and what evidence defines recovery. The decision path should remain usable when the affected network is unstable.
Auditors should sample current configurations and route observations rather than review policy documents alone. A control that exists in a standard but is bypassed by templates or emergency exceptions is not operating effectively.
Regulators should avoid reducing routing security to one technology mandate. Origin validation, filtering, change verification, monitoring, coordination and continuity address different failure modes. Evidence requirements can be technology-aware without pretending one mechanism solves all leaks.
These questions do not assume misconduct. They follow practical control. An operator with strong evidence may show that an unusual failure passed a reasonable control and was rapidly contained. An operator without evidence cannot substitute the complexity of BGP for a demonstration of due care.
The unresolved questions are part of the finding
The public evidence leaves important questions unanswered:
- What exact configuration or generated policy introduced the routes?
- Which review and validation checks ran before deployment?
- How many devices and sessions received the change?
- Which peer relationships were expected to accept or reject the announcements?
- What internal alarm first identified unexpected propagation?
- Who stopped deployment and authorized rollback?
- How were route withdrawals and forwarding recovery verified?
- Which peers changed filters after the event?
- What remediation was implemented and later exercised?
These unknowns do not erase the observed incident. They define the limit of the conclusion and the evidence needed to strengthen it.
The record supports a root-cause class: Level 3 attributed the disruption to configuration, while independent analysts observed a route leak and forwarding impairment involving AS3356. [1][2][3][4][5][6]
It supports contributing conditions: backbone scale, cross-network route acceptance, the power of more-specific prefixes and limits in external containment.
It supports a triggering event only at a high level: a routing configuration change or misconfiguration. The exact command and deployment path remain undisclosed.
It supports a detection and recovery chronology from external measurements, but not the complete internal incident timeline.
This separation matters. Root cause, contributing conditions, trigger, detection, response and recovery are related but not interchangeable. A report that calls "human error" the root cause would stop before examining why one change could escape review, propagate and require external observers to explain the failure.
The accountability standard is intended policy proven in running state
The 6 November 2017 incident was not simply a period when several websites seemed slow. It was an interdomain routing event in which externally observed BGP changes redirected traffic and degraded reachability across organizational boundaries.
The strongest public evidence is bounded. ThousandEyes reported route and forwarding changes, more-specific announcements associated with Comcast networks, increased packet loss and latency, and withdrawal by about 11:25 Pacific. [2] Level 3 attributed the disruption to a configuration issue. [1][5][6] Batfish used the event to show why policy intent should be modeled before deployment. [3] APNIC placed it among 2017's large routing-security incidents. [4]
The sources do not expose the exact internal command, every affected prefix and user, all peer policies, decision ownership or current remediation effectiveness. They do not support malicious-intent claims.
Accountability follows practical control. Level 3 controlled the change and its export policy. Peers controlled acceptance and containment. Comcast controlled subscriber communication and parts of service restoration. Customers could design alternate paths and external monitoring. End users could report impact but could not repair BGP.
The Heng.lu doctrine provides the final test. ASN, registry and policy records identify resources and preserve intent. They are accountability ledgers, not route enforcement. Running BGP announcements, peer decisions, forwarding paths and observed recovery determine whether the network actually honored that intent.
A credible repair is therefore not the statement that filters, RPKI or review procedures exist. It is evidence that a representative faulty export is rejected or contained, that unexpected route changes stop deployment, that rollback restores expected paths and that independent probes confirm user reachability.
For a backbone whose configuration can influence thousands of networks, intended policy must be proven in running state. That is the accountability standard the 2017 Level 3 event makes unavoidable.
Sources
- Wired, "How a Tiny Error Shut Off the Internet for Parts of the US"
- ThousandEyes, "Comcast Suffers Outage Due to Significant Level 3 BGP Route Leak"
- Batfish, "Don't Accidentally Break the Internet Like Level 3"
- APNIC, "14,000 Incidents: Routing Security in 2017"
- Axios, contemporaneous Comcast outage report
- KTNV, contemporaneous multi-provider outage report
- Customer Paradigm, contemporaneous customer-facing outage explanation
- Fierce Network, report distinguishing the separate 2016 Level 3 outage
- IFIP CNSM 2023, later comparative study of large-scale IP-service disruptions
- RIPEstat, AS3356 resource and routing view
- RouteViews, November 2017 BGP update archive
- RIPE NCC, Routing Information Service
- CAIDA, BGPStream
- RFC 4271, A Border Gateway Protocol 4
- RFC 7454, BGP Operations and Security
- RFC 7908, Problem Definition and Classification of BGP Route Leaks
- RFC 6811, BGP Prefix Origin Validation
- MANRS, Network Operators Actions
Member Briefing
Deeper Profile Context
Sign in with the right membership level to unlock the full briefing and source notes.
Only for Strategic Circle
Strategic Circle
Open to all readers. Unlock profile briefings after joining and signing in.
Join Strategic CircleOnly for Leadership Alliance
Leadership Alliance
For qualified IP-asset owners and management; sign in to unlock alliance briefings.
Join Leadership Alliance
