Summary

  • On 21 October 2002, a distributed denial-of-service attack materially degraded important paths to the DNS root server system. CAIDA observed abrupt round-trip-time changes from named monitoring locations, with duration varying by logical root identity. [1]
  • ICANN later described nine of the 13 logical root server addresses as swamped. That attributed formulation is evidence of attack breadth, not proof that nine complete services disappeared globally or that every resolver and user failed. [5]
  • CAIDA concluded that visible impact on global network operation was slight. Recursive caching, retry behavior and multiple logical root identities helped separate severe server and path stress from universal transaction failure. [1][20][21]
  • CAIDA's packet analysis covered E, I, K and M root links in ten-minute intervals beginning shortly after the event. Those observations are direct evidence for the monitored links, not a census of every root instance, resolver, route or application. [2][3]
  • RFC 2870 and RFC 3258 show that capacity, diverse connectivity, logging, cooperation and distributed authoritative service were recognized controls before the attack. They do not prove that every operator had deployed every control on the attack date. [8][9]
  • Later anycast, RSSAC, SSAC and large-authoritative-service guidance provide remediation and measurement criteria. They are later comparison material, not retroactive legal duties or proof of the exact 2002 topology. [6][10]-[16]
  • Responsibility was distributed among root server operators, transit and access networks, resolver operators and coordination bodies. No single institution controlled every server, route, cache, filter or user transaction.
  • The accountability standard is operational: records identify authority, while reachable routes, correct answers, resolver continuity, bounded mitigation and multi-vantage restoration evidence prove service continuity.

A record of authority is not a guarantee of service

A root-hints entry can tell a recursive resolver where a DNS root server should be found. A root-zone record can identify the authority responsible for a delegation. Neither record can make a packet cross a congested route, compel an authoritative instance to answer, preserve capacity under a flood or prove that a user’s transaction completed.

That distinction became operationally important on 21 October 2002, when a distributed denial-of-service attack sent a heavy stream of traffic toward the logical addresses of the DNS root server system. The event is often compressed into a dramatic server count. ICANN later described nine of the 13 logical root server addresses as having been “swamped.” That formulation is significant, but it is not a complete account of service availability.

It does not establish that nine complete services disappeared in every location, that every recursive resolver needed to contact a root at the same moment or that users experienced a universal failure.[5]

The accountability question is therefore more demanding than asking whether an address was under attack. It requires separating several layers that are easily conflated: traffic arriving at a server link, the response visible from a particular network path, the behavior of recursive resolvers with different cache states, and the success or failure of completed user transactions. Each layer has its own operators, evidence and failure boundaries.

The authoritative records matter because they establish which service identities resolvers are expected to use. They are essential ledgers of delegation and authority. But a ledger is not the running system. Operational continuity depends on authoritative instances that can answer correctly, routes that remain reachable, sufficient capacity, distribution across failure domains, resolvers that use cached information effectively, networks that constrain harmful traffic where they can, and operators that can coordinate mitigation and restoration.

Remove root DNS reachability from the event and there is no service-continuity question. Remove resolver caching and server stress is too easily mistaken for equivalent user harm. Remove authoritative-service distribution and the resilience design cannot be evaluated. Remove route diversity and the analysis ignores how demand reaches healthy capacity. Remove source-address filtering and responsibility at the traffic-origin edge disappears. Remove measurement vantage points and local observations become unjustified global claims.

Remove operator coordination and there is no credible account of how a distributed service detects, mitigates and declares recovery.

The 2002 attack made those dependencies visible. It did not show that one institution possessed exclusive control. It showed that the service users experience is assembled from records, running code, routing, caches and autonomous operational decisions. Accountability must follow those practical controls and the evidence each controller can produce.

What was observed on 21 October 2002

CAIDA’s contemporary measurement account reported abrupt round-trip-time degradation at approximately 22:00 UTC. From the UCSD vantage point, all monitored roots except I and M showed a performance change, but the duration was not uniform. CAIDA reported effects lasting roughly an hour for F, G and L, about five to ten minutes for A and B, and somewhat more than ten minutes for J.[1]

Those observations provide evidence of a serious event. They do not provide a universal map of root service from every network. The measurements depended on monitors in San Diego and San Jose and therefore described service as visible across particular paths from particular locations. A root address that responded poorly from one monitor might have presented a different condition to a resolver using another upstream, route or geographic path. Conversely, a root that looked responsive from a measurement site might still have been difficult to reach from another network.

This limitation is not a weakness unique to CAIDA. It is a basic property of distributed-service measurement. A monitor records what its path allows it to see. Its results become broadly meaningful only when the location, metric, time interval and target are stated alongside the conclusion.

CAIDA also concluded that the visible impact on global network operation was slight.[1] That finding constrains any retelling of the attack. It rules out treating severe traffic and root-server performance changes as automatic proof of universal user failure. It also requires an explanation: the DNS architecture contains buffers between an authoritative root address and an individual transaction, most notably recursive caching and the availability of multiple root service identities.

D-Root’s operator history independently identifies 21 October 2002 as the date of a massive attack and notes that operators subsequently prepared an analysis of the event.[4] That record corroborates the date and the operational recognition of the incident. It does not, by itself, establish identical conditions at every root identity or every physical instance.

CAIDA’s later packet analysis supplies another bounded view. It examined traffic collected from links serving the E, I, K and M roots, beginning shortly after the attack, and grouped observations into ten-minute intervals. Its distributions of requests and observed clients are direct measurements from those links.[2] The broader CAIDA root-traffic work explains the dataset and methodology in which those observations sit.[3]

The E, I, K and M link data must not be treated as a census of every root identity, physical instance, resolver, access network or application. It began after the attack was underway, covered named links and used defined observation intervals. It is valuable precisely because its boundary is knowable. The correct use of that evidence is to state what appeared at the monitored links and when, then resist turning those samples into unsupported totals for the entire system.

Three source-specific descriptions can therefore coexist without contradiction:

  • CAIDA observed path-performance changes whose duration differed by root identity from its monitoring locations.[1]
  • CAIDA later analyzed packets and apparent clients at selected E, I, K and M links in ten-minute intervals.[2]
  • ICANN’s 2007 comparison described nine of 13 logical root server addresses as swamped during the 2002 attack.[5]

They observe different objects. One concerns measured path performance, another concerns packets at selected links, and the third is a later institutional summary framed around logical addresses. Responsible analysis preserves those distinctions.

Why “nine of 13” is the beginning of the inquiry

The number thirteen refers to logical root server identities. A logical identity is not necessarily equivalent to one machine, one site or one route. Its operational realization depends on how the responsible operator deploys authoritative capacity and advertises the service address.

That point matters even though the physical and topological distribution in 2002 was less extensive than the root system’s later deployment. The frozen public evidence does not enumerate every physical instance, active route or local mitigation arrangement on the day of the attack. It therefore cannot support a precise reconstruction of which hardware or sites were simultaneously reachable from every part of the Internet.

ICANN’s formulation should be preserved in its attributed form: nine of the 13 logical root server addresses were described as swamped.[5] “Swamped” communicates that the attack overwhelmed important service paths or capacity. It does not define a universal outage boundary. It does not say that all authoritative capacity associated with each address failed everywhere. It does not reveal every resolver’s retry decision or cache state. It does not count completed transactions.

A server-address count omits at least four dimensions.

First, it omits vantage point. Reachability is relational: a resolver reaches an address through a specific route from a specific network. A response failure observed in California is evidence about that path and time, not a direct observation from Europe, Africa, Asia or another North American network.

Second, it omits time. CAIDA’s reported performance changes did not all last equally long.[1] A count without a time interval can combine brief degradation at one address with a longer condition at another and make them appear operationally identical.

Third, it omits service distribution. One logical identity can be represented by more than one service location when a distributed addressing design is deployed. The existence, extent and behavior of such distribution must be demonstrated for the date and identity being discussed. It cannot be inferred from later topology.

Fourth, it omits the resolver layer. A recursive resolver may answer using valid cached referrals without sending a root query for each user request. Another resolver may need fresh information, retry a different root identity or encounter a different path. The server-address headline cannot resolve those differences.

The count is consequently evidence of attack severity, not a self-sufficient measure of user harm. The right question is not whether the number should be minimized. It is what the number measured, what it omitted and what additional evidence is needed to translate it into a service conclusion.

This distinction protects accountability rather than diluting it. If every overloaded address is casually equated with universal service disappearance, operators cannot tell which controls actually contained harm. Caching, alternate authoritative capacity, route diversity and coordination become invisible. If the count is dismissed because many transactions still succeeded, the attack’s pressure on critical infrastructure is understated. A defensible account must hold both facts at once: the attack materially degraded important root service paths, while the available evidence does not establish universal failure.

Resolver caching separates infrastructure stress from user harm

The DNS resolution process does not require every user query to travel to a root server. Recursive resolvers retain DNS information for its permitted cache lifetime. When a resolver already holds the referral needed to continue resolution, it can use that unexpired information without contacting a root server for that transaction.[20], [21]

Caching therefore changes the relationship between an attack against authoritative infrastructure and the experience visible to users. It creates a temporary service buffer. The buffer is neither unlimited nor uniform.

Two resolvers can face the same attack and experience different outcomes because their caches contain different records with different remaining lifetimes. One may already possess the referral required for a popular domain. Another may need to ask the root for information it has not cached or whose cached copy has expired. Their upstream routes may also differ. Even if they contact the same logical root address, they may not observe the same response condition.

This explains why a severe attack on root server addresses need not produce an equally severe and simultaneous failure across applications. It also explains why server telemetry alone cannot settle the user-impact question. A root operator can demonstrate that traffic surged or that responses degraded at an instance. That evidence is essential, but it does not show which recursive resolvers needed the affected service during the interval or which user transactions failed.

The reverse inference is also unsafe. Limited immediate user-visible harm does not mean the infrastructure event was inconsequential. Cached data ages. A prolonged loss of reachable authoritative service would progressively expose resolvers that require information no longer available locally. The architecture can absorb a shock without making the shock irrelevant.

Caching belongs in the accountability map because recursive resolver operators control important aspects of this buffer. They manage cache behavior, retries, root hints and user-facing monitoring. Their evidence can show whether resolvers continued answering, shifted queries among root identities, encountered timeouts or exhausted useful cached information. Without resolver-side data, an assessment remains trapped between server distress and anecdotal user experience.

The root server system and recursive resolver population also operate on different clocks. A root instance can experience an immediate traffic spike. A monitor can observe increased round-trip time seconds or minutes later. Resolver effects depend on cache state and retry behavior. A user transaction adds application-specific timing and tolerance. Combining those clocks into one statement such as “the DNS was down” erases the causal chain.

The bounded conclusion reported by CAIDA—that the visible impact on global network operation was slight—fits this layered model.[1] It should not be generalized into a claim that no users were affected, because exact user-level failure rates are not available. Nor should it be discarded in favor of a dramatic address count. It is evidence that the wider system, including caches and remaining authoritative reachability, continued to deliver substantial service despite intense pressure.

Accountability thus requires both stress evidence and continuity evidence. The former shows where infrastructure came under load. The latter shows whether resolvers and transactions continued to obtain timely, correct answers. Neither substitutes for the other.

Distributed operation changes the ownership of resilience

The DNS root server system is distributed not only in topology but also in operational authority. Individual root server operators control their own instances, capacity planning, upstream connectivity, local filtering, monitoring and incident response. Transit and access networks control other parts of the path. Recursive resolver operators control caches and retries. Coordination and advisory bodies connect those domains but do not erase their autonomy.

This means responsibility cannot be reduced to the organization associated with the root zone or to the institution that later published a factsheet. ICANN, the IANA functions and RSSAC-related structures have root-zone, coordination and advisory roles. They did not constitute a single command system operating every root instance during the 2002 attack.

The distinction between a recordkeeper and an operator is central. Root-zone and root-hints records identify which logical services hold authority. Accuracy in those records is indispensable: incorrect authority information would direct resolvers toward the wrong service. But correct records cannot ensure that routes are available, that an instance has capacity or that an upstream network filters harmful traffic.

Operational continuity is produced by the parties that control those layers. A root operator can add capacity, distribute service, diversify upstreams and collect instance telemetry. A transit provider can provision paths, manage congestion and enforce source-validity controls within its domain. An access network can constrain implausible source traffic leaving customer networks. A resolver operator can maintain reliable cache and retry behavior. Coordination bodies can establish shared expectations and evidence formats. No one role can replace all the others.

This distribution prevents a simple assignment of exclusive blame, but it does not eliminate accountability. It makes accountability more precise. Each operator should be evaluated against the controls it actually possessed, the evidence available at the time and the actions it could reasonably take without assuming powers held elsewhere.

For root operators, relevant questions include whether legitimate queries could reach usable authoritative capacity, whether connectivity was diverse enough to avoid a single path bottleneck, whether monitoring distinguished an instance problem from a broader route problem and whether restoration could be demonstrated from outside the operator’s own network.

For transit and access networks, the questions concern path capacity, routing behavior and controls at the network edge. A root operator cannot directly validate source addresses on every originating access network. An access provider cannot dictate how every root service distributes capacity. Their duties are different because their controls are different.

For resolver operators, the questions concern continued service to users, cache behavior, retry outcomes and the ability to distinguish upstream root distress from local resolver or access failures.

For coordination bodies, accountability concerns the quality of common expectations, information exchange, incident analysis and comparable metrics—not a fictional power to issue instantaneous commands to every independently operated service.

Users occupy the least empowered position. They can observe slow or failed transactions but ordinarily cannot inspect root-link load, routing changes, cache contents or operator coordination. A system that places the burden of proof on users would invert the control structure. Evidence should come from the operators that possess the relevant telemetry.

Distributed responsibility is therefore neither centralized sovereignty nor operational ambiguity. It is a control map. The map makes it possible to ask who could observe a condition, who could change it, who depended on another party and what evidence should survive for later review.

Capacity and diversity were recognized controls before the attack

The 2002 event did not occur in a conceptual vacuum. RFC 2870, published before the attack, set out operational requirements for root name servers. It addressed standards-compliant service, capacity above measured peak demand, diverse connectivity, authoritative-only operation, logging and cooperation in security analysis.[8]

Those expectations are directly relevant to the attack because they identify the kinds of controls that make authoritative infrastructure more resilient. Capacity headroom can absorb abnormal demand up to a point. Diverse connectivity can reduce dependence on one provider or path. Authoritative-only operation narrows the service role. Logging and cooperative analysis support detection and reconstruction.

The document cannot, however, prove the deployment state of every operator on 21 October 2002. An operational requirement in an RFC is not evidence that every root identity had implemented the same architecture, possessed equivalent capacity or recorded the same telemetry. Nor does it establish a legal duty, negligence or breach. Those conclusions would require evidence beyond the technical record supplied here.

RFC 2870 is best used as pre-event context. It demonstrates that capacity, connectivity diversity, logging and cooperation were already understood as material operational controls.[8] It permits an accountability inquiry into those areas without pretending that later architecture was already universal. It does not answer the inquiry by itself.

RFC 3258, published in April 2002, described using shared unicast addresses to distribute authoritative name service.[9] The design allows a service address to be originated from multiple locations, with routing determining which location receives a query. It also introduces its own operational boundaries: placement matters, routing behavior matters, and authoritative data must remain consistent across the distributed service.

The timing is important but must be handled carefully. RFC 3258 shows that distributed authoritative service through routing was documented before the October attack. It does not establish that every root identity had deployed it, that all deployments were equivalent or that the root system already possessed its later anycast footprint.

Capacity and distribution also solve different problems. More capacity at one location may withstand a larger local flood, but it remains dependent on the routes and upstream links serving that location. More locations can spread demand and reduce shared failure, but distribution is only useful when routes direct legitimate queries to reachable capacity and the instances provide consistent answers. Route diversity without sufficient service capacity can merely expose more overloaded paths. Service capacity without route diversity can remain unreachable.

The attack consequently tests a chain rather than a single control:

  1. The authority record must identify the correct logical service.
  2. Routing must deliver queries to an operational instance.
  3. The path and instance must have sufficient usable capacity.
  4. Distributed instances must return consistent authoritative answers.
  5. Resolver behavior must make effective use of the available service.
  6. Monitoring must reveal which link in the chain is impaired.
  7. Operators must coordinate when the impaired link crosses organizational boundaries.

Every link has an evidence requirement. A configuration file can prove the intended authority. A route observation can show advertised reachability. Instance telemetry can show load and response behavior. Resolver measurements can show practical resolution. Transaction tests can show user-facing completion. A credible post-incident account should not use one category as a substitute for all the others.

Shared unicast and the danger of rewriting attack-day topology

Later resilience improvements can make an earlier system appear simpler than it was. The root server system’s subsequent expansion through anycast is especially prone to this distortion.

RFC 4786 later defined an anycast operational model in which the same service address is advertised from multiple discrete locations.[10] RFC 7094 developed further architectural considerations for anycast distribution, including the relationship among topology, routing and service behavior.[11] RFC 7720 later described protocol and deployment requirements for root name service.[12]

These publications provide useful comparison criteria. They explain why a logical service address should not automatically be interpreted as one physical machine and why routing is part of service delivery. They also help identify operational risks: placement can be uneven, routes can shift demand, sites can have different capacity, and distributed instances require consistent service behavior.

They are not a license to project later deployment backward. The evidence supports neither a claim of universal root anycast on 21 October 2002 nor a complete identity-by-identity map of attack-day distribution. RFC 3258’s pre-event publication establishes that shared-unicast distribution was a documented technique.[9] It does not establish universal implementation.

ICANN’s 2007 factsheet compared the later root attack with the 2002 event and attributed the lower user impact of the later attack partly to anycast deployment and improved operator coordination developed after 2002.[5] That comparison supports the conclusion that distribution and coordination became more significant resilience controls. It does not convert later controls into retroactive duties or prove that one remediation accounted for every difference between the events.

The responsible comparison is causal and limited. Distributed service can make it harder for a flood aimed at one logical address to consume all associated capacity because routing may deliver traffic to multiple locations. Route and upstream diversity can isolate some failures. More observation points can reveal regional differences. Coordination can help operators exchange attack indicators and protect legitimate traffic.

But anycast is not an incantation. The same address at multiple sites does not guarantee equal reachability, balanced load, independent upstreams or sufficient capacity. Routing policies determine where traffic goes. A badly placed or underprovisioned site can still suffer. A route change can redirect demand. Monitoring that aggregates all instances under one logical label can hide local distress.

The accountability value of distribution therefore lies in demonstrable outcomes. Operators should be able to show which instances served an address, which routes exposed them, how traffic shifted, whether legitimate responses remained timely and correct, and whether failures remained contained. The existence of an anycast label is not enough.

This returns the analysis to running service. A logical address listed in root hints identifies where service should be available. A distributed deployment creates more ways to realize that identity. Only measurement can show whether the realization worked from diverse networks during an attack.

Measurement must identify its clock, layer and vantage point

The 2002 evidence demonstrates why infrastructure accountability needs disciplined measurement language. A statement about “the root” can refer to at least five different objects:

Measurement layer What it can establish What it cannot establish alone
Link and packet load Traffic volume and apparent clients visible at a monitored link during a defined interval Conditions at every root link or the success of user transactions
Path performance Reachability or response degradation from a named monitor to a logical address Reachability from every resolver or region
Authoritative service Whether an observed instance returned timely, correct DNS answers The cache state and behavior of downstream resolvers
Recursive resolver Whether resolution continued through caches, retries and available roots Conditions experienced by every application or user
Completed transaction Whether a specific user-facing operation succeeded from a defined network The global health of every underlying DNS component

CAIDA’s event account sits primarily in the path-performance layer. It reported changes in round-trip time around 22:00 UTC and different apparent durations among monitored roots.[1] Those measurements are strong evidence when expressed as observations from the stated sites. They become weaker if transformed into global availability claims.

The later E, I, K and M analysis sits at the link and packet layer. Its ten-minute groupings provide a clock and its monitored links provide a scope.[2] The dataset context explains how such root-traffic observations were assembled.[3] The analysis can characterize what those links saw. It cannot establish every physical origin, every route or every resolver outcome.

ICANN’s “nine of 13” account is a logical-address summary.[5] It captures the breadth of pressure across named services but does not replace path, instance, resolver or transaction evidence.

A credible incident reconstruction should therefore attach four qualifiers to every major statement:

  • Object: Was the observation about a logical address, a physical or topological instance, a link, a route, a resolver or a transaction?
  • Vantage point: From which network or monitoring position was it observed?
  • Metric: Was the evidence traffic volume, round-trip time, response rate, correctness, timeout behavior or completed resolution?
  • Interval: When did the condition begin, how was duration measured, and when was restoration confirmed?

Without those qualifiers, different measurements can be made to contradict each other when they actually describe different layers. A root link can be heavily loaded while a resolver continues answering from cache. A monitor can see degraded response from one route while another route remains usable. A transaction can succeed even though one attempted authoritative query timed out and a retry reached another service identity.

Later RSSAC publications offer comparison criteria for making root-service expectations and measurements more consistent. RSSAC’s service-expectation work frames the root server system in terms of the service delivered, while its common measurement framework seeks comparable evidence across operators.[13], [14] RFC 9199 likewise provides later operational considerations for large authoritative DNS server systems.[16]

These later documents should not be presented as obligations that governed every operator in 2002. Their value is retrospective and forward-looking: they show how evidence can be structured so that a future event is easier to declare, compare and close.

Measurement accountability also applies to restoration. An operator’s internal graph returning to normal is useful but not sufficient if external resolvers still cannot obtain answers. A monitor recovering does not prove recovery everywhere. A resolver succeeding once does not establish sustained stability. Restoration should be supported by several layers: instance health, route reachability, authoritative correctness, resolver success and geographically or topologically diverse external observations.

The objective is not an impossible global census. Distributed systems rarely provide one. The objective is bounded evidence whose scope is explicit enough that decision-makers can tell what is known, what is inferred and what remains unknown.

Practical control determines practical accountability

A distributed service needs a responsibility model that follows actual control. That model can be stated without alleging fault.

Actor Principal controls Evidence expected
Root server operators Instance placement, authoritative capacity, upstream diversity, local filtering, monitoring and incident response Per-instance and per-link health, response behavior, route context, mitigation timing and restoration evidence
Transit and peering networks Path capacity, route propagation, congestion management and network-domain filtering Route and traffic changes, affected paths, filtering actions and reachability from relevant networks
Access networks Customer-edge controls and source-address plausibility within their domains Deployment scope, exceptions, validation results and attack-related observations where available
Recursive resolver operators Cache behavior, retries, root hints and user-facing resolution monitoring Cache-dependent success, timeout and retry patterns, root-selection outcomes and resolver restoration
Coordination and advisory bodies Shared expectations, information exchange, evidence formats and post-incident analysis Timely notices, common terminology, comparable measurements, recorded decisions and bounded findings
End users Application requests and local observations Transaction symptoms, without an expectation that users reconstruct hidden infrastructure state

Root operators have the most direct control over authoritative service but not over every packet path. They can distribute capacity, select upstreams, monitor instances and apply local mitigations. They cannot alone prevent every originating network from emitting harmful traffic.

Transit and peering networks control whether traffic can reach authoritative capacity across particular paths. They can influence congestion, route availability and the movement of demand among sites. The supplied evidence does not reconstruct every route or peering decision during the 2002 attack, so no specific route intervention should be inferred. The absence of a complete route record is itself an accountability lesson: service claims should be accompanied by enough routing evidence to distinguish instance exhaustion from path failure.

Access networks possess a different control. They are positioned to assess whether traffic leaving their domain uses source addresses plausible for that domain. That does not make them controllers of the root system. It makes them responsible for a risk boundary that attacked servers cannot fully enforce at the destination.

Recursive resolver operators control the component closest to ordinary DNS use. Caching can preserve continuity, retries can locate remaining service, and monitoring can reveal whether root stress is translating into resolution failure. A resolver operator does not control root capacity, but it can provide decisive evidence about user-impact propagation.

ICANN and related coordination structures occupy another layer. Root-zone and institutional roles make ICANN relevant to the service ecosystem and to later analysis, but they do not amount to exclusive operational control over independently run root services. The same is true of advisory structures: they can define expectations, facilitate coordination and improve evidence without directly operating every instance.

The SSAC advisory on DDoS risk and coordinated DNS controls, along with the institutional record responding to that work, illustrates the later development of this coordination layer.[6], [7] The root server operator threat-mitigation framework similarly reflects a distributed model in which resilience arises from operator controls and cooperation rather than a single command point.[15]

Control ownership should be evaluated through three questions.

First, who could observe the condition? A root operator sees instance and link telemetry. A transit network sees traffic and routes in its domain. A resolver operator sees timeouts, cache use and retries. A user sees transaction symptoms.

Second, who could change the condition? The root operator can add or redistribute service capacity. The upstream can alter routing or mitigation. The access network can constrain implausible source traffic. The resolver operator can maintain robust retry and caching behavior. A coordination body can align communications but cannot substitute for those operational actions.

Third, who can prove restoration? No single actor has every piece. Root operators can show service recovery, networks can show route and traffic normalization, resolver operators can show renewed resolution success, and external monitors can test reachability. A credible closeout joins these records without pretending they came from one controller.

This model turns accountability into an engineering discipline. It avoids both extremes: blaming one institution for a system it did not exclusively operate, and treating distributed operation as a reason that nobody must explain outcomes.

Ingress filtering is an upstream responsibility, not a universal cure

Source-address filtering belongs in the analysis because denial-of-service traffic can exploit weaknesses far from the attacked service. RFC 2827 describes ingress filtering intended to reduce traffic carrying forged source addresses.[17] RFC 3704 develops filtering considerations, including complications created by multihomed networks and asymmetric routing.[18] RFC 4732 treats denial of service as an Internet-wide engineering problem requiring attention across multiple parts of the network.[19]

These documents identify a control boundary. A network that knows which source addresses should legitimately originate from its customers or downstreams is better positioned to reject implausible traffic than a root server receiving packets after they have crossed multiple networks.

That principle does not establish that the 2002 attack depended on a particular spoofing method. The supplied public evidence does not identify the complete source population, attacker, motive or packet-generation method. It would be improper to infer those facts from the existence of anti-spoofing standards.

Ingress filtering is also not a single-network cure for a distributed flood. Filtering forged sources at one access network does not prevent harmful traffic from other networks, nor does it stop traffic using valid source addresses. Its effectiveness depends on deployment across relevant origin edges, accurate policy and accommodation of legitimate routing complexity.

The control remains important because destination-side defenses cannot repair every weakness at the origin edge. If forged-source traffic is permitted to leave an access network, the attacked service sees a symptom after the more discriminating enforcement point has been passed. Conversely, overly broad filtering can harm legitimate multihomed traffic. Accountability requires evidence that controls are both effective against implausible sources and precise enough to preserve legitimate connectivity.

The appropriate questions are therefore bounded:

  • Did a network deploy source-validity controls within the domain it could actually govern?
  • Were exceptions for multihoming and asymmetric paths understood and tested?
  • Did operators collect evidence showing what the controls accepted or rejected?
  • Could filtering changes be correlated with improvements in legitimate service?
  • Did mitigation shift harmful traffic elsewhere or create new reachability failures?

None of those questions identifies the 2002 attacker. None assigns exclusive responsibility to access networks. They ensure that the analysis does not place the full burden on authoritative servers when some relevant controls exist upstream.

Route diversity, peering and transit capacity belong beside filtering. A root instance may have sufficient computing capacity yet remain inaccessible if its upstream path is saturated. Another instance may be healthy but receive little traffic because routing does not direct affected resolvers toward it. Filtering reduces certain traffic risks; diversity preserves alternate delivery paths; service distribution creates additional capacity endpoints. The controls complement one another but are not interchangeable.

Coordination is an operational control, not a claim of central command

A distributed operator model depends on coordination precisely because no single party controls the entire system. During a fast-moving attack, operators need common terms for what they are observing, channels for exchanging bounded evidence and a way to distinguish local mitigation from system-wide restoration.

Coordination should not be confused with permission. A root operator must be able to protect its own service without waiting for a central institution to direct every technical action. A transit network must act within its domain. A resolver operator must preserve local service. Coordination becomes valuable when those autonomous actions affect shared outcomes.

The 2002 records show why common evidence matters. CAIDA described path performance from named monitors.[1] Its later analysis described packets at selected root links.[2] ICANN summarized logical addresses.[5] D-Root recorded the event in operator history.[4] Each account is useful, but their different objects must be reconciled carefully.

Later SSAC, RSSAC and operator publications can be read as responses to that evidence problem. They emphasize coordinated controls, service expectations, shared measurements and threat mitigation.[6], [13]-[15] Their relevance lies in improving future observability and response. They should not be used to declare that all such mechanisms were mandatory or deployed in October 2002.

Effective coordination has measurable outputs. Operators can timestamp when they recognized a shared event, state which service identities or instances were affected, record routing and filtering changes, identify the evidence used to declare recovery and preserve disagreements about scope. A coordination process that produces only a global label—“up” or “down”—does not capture a distributed system.

The public conclusion should remain narrower than internal certainty. If operators possess incomplete visibility, the correct declaration is a bounded one: service recovered at specified instances and external vantage points, while other regions remain unverified. Explicit uncertainty is more accountable than unsupported universality.

A measurable resilience and restoration test

The central lesson of the attack is not that one later technology solved root DNS risk. It is that resilience claims must be converted into evidence across the whole service path.

A defensible accountability test can be organized into seven connected stages.

1. Authority integrity

The first question is whether resolvers have accurate records identifying the expected root service identities. Root-zone and root-hints information performs this ledger function. If those records are wrong, running capacity at the correct destination may never be reached.

Authority integrity is necessary but not sufficient. Passing this stage proves that the system points toward the intended services. It does not prove those services are reachable.

2. Reachability from diverse networks

The second stage tests whether the logical addresses can be reached from multiple independent network locations. Measurements should name their vantage points, routes where available, metrics and intervals.

CAIDA’s observations demonstrate why this discipline matters. The reported round-trip-time changes were real measurements from particular monitors.[1] A modern resilience assessment should expand the number and diversity of such perspectives, but it must still resist claiming more geography than the monitors cover.

A pass at this stage does not require every probe to report identical performance. It requires operators to understand where service is healthy, degraded or unverified and to avoid hiding regional failures inside a global average.

3. Usable authoritative service

Reachability must lead to timely, correct authoritative answers. A route to an address that returns no usable response does not deliver continuity. Operators should distinguish link saturation, instance exhaustion, incorrect answers and local mitigation that drops legitimate traffic.

Distribution should be described at the instance level where disclosure is operationally safe. The relevant evidence includes which service locations remained available, whether capacity was independent enough to avoid a common bottleneck and whether the same authoritative data was delivered across them.

RFC 2870 provides the pre-event context for capacity, diverse connectivity, logging and cooperation.[8] RFC 3258 supplies the early distributed-service model and its routing and consistency concerns.[9] Later anycast and root-service documents refine the comparison.[10]-[12] None of them substitutes for attack-specific telemetry.

4. Resolver continuity

The fourth stage tests whether recursive resolvers can continue obtaining answers through cache, retries and reachable root identities. This stage prevents the assessment from equating server distress with transaction failure.

Resolver results should be separated by cache condition where possible. A warm cache demonstrates that the architecture’s buffer worked. A query requiring uncached or expired information tests current authoritative reachability more directly. Both matter, but they answer different questions.

RFC 1034 and RFC 1035 provide the foundation for the resolver and caching behavior that creates this boundary.[20], [21] The exact cache state of the global resolver population during the 2002 event remains unknown, so it cannot be reconstructed from server data alone.

5. Legitimate transaction completion

The fifth stage measures whether real or representative user-facing resolutions complete. Transaction evidence should name the access network and time window rather than presenting a few successful tests as global proof.

This stage is where infrastructure performance becomes user impact. It should still be interpreted cautiously. A transaction can fail because of a local resolver, access path, authoritative dependency below the root or application timeout. Root-related attribution requires evidence connecting the failure to the relevant root-service condition.

CAIDA’s conclusion of slight visible global operational impact is an important event-specific boundary.[1] It does not quantify every user experience, but it prevents a claim of universal collapse.

6. Traffic-origin and route controls

The sixth stage tests controls outside the authoritative service. Networks should be able to explain their source-address validation posture, path capacity and relevant routing or filtering changes.

RFC 2827 and RFC 3704 identify ingress-filtering principles and their operational limits.[17], [18] RFC 4732 places denial-of-service mitigation in a distributed engineering context.[19] Evidence at this stage should not assume that all attack traffic used forged sources. It should show which controls were available, where they applied and whether changes preserved legitimate traffic.

Route diversity must also be tested rather than asserted. Multiple upstream names do not prove independent failure domains. Multiple routes do not guarantee that affected resolvers will reach healthy capacity. A useful assessment connects route observations to service results.

7. Coordinated declaration and restoration

The final stage asks whether operators can state when the event was declared, which service boundaries were affected, what changed and how recovery was verified.

Later RSSAC measurement work offers a framework for comparable root-system evidence, while operator threat-mitigation and large-authoritative-service guidance provide additional comparison points.[13]-[16] These later materials should guide current expectations without being misrepresented as retroactive attack-day obligations.

Restoration should require agreement among several indicators:

  • Traffic and response conditions at affected instances have stabilized.
  • Routes make healthy service reachable from diverse external networks.
  • Authoritative answers remain correct and timely.
  • Resolver tests succeed under relevant cache conditions.
  • Legitimate transaction tests recover.
  • Mitigation does not create an equivalent accessibility failure.
  • Remaining unknown regions or service identities are explicitly recorded.

No single indicator can prove the entire chain. Together, they can support a bounded, reproducible declaration.

This framework makes remediation measurable. “Add anycast,” “increase capacity” or “improve coordination” are incomplete promises. A remediation should state which failure boundary it addresses and what evidence will show that it worked. Additional instances should improve reachability or failure isolation. More capacity should increase usable headroom at defined paths. Filtering should reduce harmful traffic without excluding legitimate sources. Coordination should shorten detection, align scope and produce comparable restoration evidence.

The attack exposed a continuity problem, not a sovereignty problem

The 2002 event can be misunderstood as a contest over which institution controlled the root. That framing misses the operational failure boundary.

The root records identified the logical services expected to answer. The attack did not primarily challenge the existence of those records. It challenged whether resolvers could reach running authoritative capacity through the available network.

That distinction matters because records and service have different accountability properties. A record can be audited for accuracy and authorized change. A service must be tested for reachability, correctness, capacity and continuity. The institution maintaining or coordinating a record does not automatically control every route and server that realizes it.

The root server operators’ practical autonomy is therefore not an obstacle to accountability. It is a fact that the accountability design must reflect. Each operator should be able to produce evidence for its own instances and collaborate on a system-level view. Transit and access networks should account for their paths and edge controls. Resolver operators should account for the continuity presented to users. Coordination bodies should preserve the common record without pretending to command every operational action.

This reality-layer approach is stricter than a governance slogan. It asks whether the system ran, where it was reachable, which controls absorbed the attack and how recovery was demonstrated. It also prevents authority records from being treated as magical guarantees. A correct name in a ledger does not move packets.

The same approach constrains claims about prevention. No supplied evidence shows that one operator or institution could have prevented the attack alone. More capacity at one root would not control traffic at other roots. Filtering at one access network would not constrain all sources. Caching at resolvers would not preserve data indefinitely. Coordination would not create capacity by itself. Resilience emerged from the combined effect of multiple controls.

The accountability test is consequently plural but not vague. It asks each controller for evidence at the layer it operates and asks the system as a whole to demonstrate continuity across those layers.

What the public evidence cannot establish

Several important facts remain unknown from the supplied record.

The attacker’s identity and motive are not established. The complete population of systems or sources involved is not known. The evidence does not justify attributing the attack to a named person, organization or class of actor.

Exact packet rates at every root identity and physical instance are not available. CAIDA’s packet work covered E, I, K and M links beginning shortly after the attack and organized observations into ten-minute intervals.[2] Those data should not be extended to unmonitored links.

Every attack-day route and mitigation change is also unknown. The public evidence does not supply a complete BGP, peering or transit reconstruction for every root operator. It cannot support claims that a particular route decision caused or ended the event unless independently documented.

The full physical topology active on 21 October 2002 is not enumerated. Later anycast expansion cannot fill that gap. It would be inaccurate to describe the later distribution model as universally present during the attack.

Resolver cache states and exact user-visible failure rates are unavailable. CAIDA’s slight-impact conclusion is an important bound, but it is not a census of every resolver or user.[1] Some transactions may have failed; the record does not quantify them universally.

Complete internal operator timelines, coordination records, costs and legal allocations are also outside the evidence. Technical standards and later advisory documents identify engineering controls, but they do not establish negligence, illegality, breach or legal liability. Those findings would require facts and legal analysis not present here.

The effectiveness of every post-event remediation likewise cannot be assumed. ICANN’s later comparison associated reduced user impact in the 2007 attack partly with anycast deployment and operator coordination after 2002.[5] That supports a limited comparison, not a universal claim that every later control worked equally well under every condition.

These unknowns should remain visible. Precision about uncertainty is part of infrastructure accountability because it prevents a local observation, institutional summary or later design improvement from being converted into an unsupported historical certainty.

The essential accountability lesson

The 2002 root DNS attack was serious because it placed coordinated pressure on infrastructure that recursive resolvers depend on to navigate the DNS hierarchy. Its significance does not require a claim that the Internet nearly stopped or that every user lost service.

CAIDA observed abrupt performance degradation around 22:00 UTC and reported different durations among monitored root identities from its vantage points.[1] Its later packet analysis documented traffic at selected E, I, K and M links.[2] D-Root’s history corroborated the event’s operational significance.[4] ICANN later summarized the breadth of the attack as nine of 13 logical root server addresses being swamped.[5] CAIDA nevertheless concluded that visible global operational impact was slight.[1]

Those facts fit together when the system is examined as a chain. Logical addresses identify services; they are not identical to complete physical deployments. Routes determine which capacity a resolver can reach. Distributed authoritative instances create alternatives but depend on topology and consistency. Recursive caching reduces immediate dependence on live root queries. Transit and access networks control parts of the traffic path and source-validity boundary. Measurement determines whether conclusions are local, regional or system-wide. Coordination connects autonomous operators during mitigation and restoration.

The attack therefore made resilience an accountability test. It required more than evidence that records remained intact or that some server answered somewhere. It required proof that correct authoritative service remained reachable enough, from enough paths, for resolver and transaction continuity to be sustained.

Later anycast, RSSAC measurement and threat-mitigation work can be evaluated as responses to that test.[10]-[16] They should not be turned into attack-day mythology or retroactive legal duties. Their value is that they make control and evidence more explicit.

The durable principle is simple: records identify authority; running, reachable service proves continuity. A resilient distributed service must be able to show both. Its operators must demonstrate what they controlled, what they observed, how they coordinated and how they knew recovery was real. Anything less leaves an accurate ledger pointing toward a service whose availability cannot be proved.

Sources

  1. https://www.caida.org/projects/dns/oct02dos/
  2. https://www.caida.org/catalog/papers/2010_understanding_dns_evolution/roottraffic/2002-analysis/2002-10-21/
  3. https://www.caida.org/catalog/papers/2010_understanding_dns_evolution/roottraffic/
  4. https://d.root-servers.org/history.html
  5. https://www.icann.org/en/system/files/files/factsheet-dns-attack-08mar07-en.pdf
  6. https://www.icann.org/en/groups/ssac/dns-ddos-advisory-31mar06-en.pdf
  7. https://archive.icann.org/historical-resolution-tracking-feature/2006-03-31-ssac-report-dns-distributed-denial-service-ddos-attacks-tld-and-root-name-system.html
  8. https://www.rfc-editor.org/rfc/rfc2870
  9. https://www.rfc-editor.org/rfc/rfc3258
  10. https://www.rfc-editor.org/rfc/rfc4786
  11. https://www.rfc-editor.org/rfc/rfc7094
  12. https://www.rfc-editor.org/rfc/rfc7720
  13. https://www.icann.org/en/system/files/files/rssac-001-draft-02may13-en.pdf
  14. https://itp.cdn.icann.org/en/files/root-server-system-advisory-committee-rssac-publications/rssac-002-20nov14-en.pdf
  15. https://root-servers.org/media/news/Threat_Mitigation_For_the_Root_Server_System.pdf
  16. https://www.rfc-editor.org/rfc/rfc9199
  17. https://www.rfc-editor.org/rfc/rfc2827
  18. https://www.rfc-editor.org/rfc/rfc3704
  19. https://www.rfc-editor.org/rfc/rfc4732
  20. https://www.rfc-editor.org/rfc/rfc1034
  21. https://www.rfc-editor.org/rfc/rfc1035