Summary

  • Frozen incident boundary: This article covers the AS9121/TTNet routing-table leak observed on 24 December 2004. It does not merge that event with later Turkish connectivity incidents, unrelated route hijacks, or other leaks involving different autonomous systems.
  • Scale must be attributed: The NANOG reconstruction and later research describe a leak involving more than 100,000 prefixes, representing a large majority of the then-visible global routing table from some measurement points. The exact count depends on collector, interval, duplicate handling, and event definition. It is evidence, not a universal operator log.
  • The failure crossed relationship boundaries: Routes learned in one context were advertised where they were not intended to go. Other networks accepted and propagated them under their own local policy. The result was not merely a TTNet service outage; it was a distributed change in global route state.
  • Responsibility follows control: AS9121 controlled its import, export, route generation, deployment, monitoring, and withdrawal. Direct providers and peers controlled customer-prefix filters, AS-path constraints, maximum-prefix limits, alerting, and onward export. Further importers controlled their own acceptance and containment.
  • A registry is evidence, not enforcement: ASN, address, IRR, and later RPKI records can describe expected origins and resource holders. They do not configure router policy by themselves. Running code and loaded policy determine which route is accepted and propagated.
  • Modern controls are complementary: Explicit eBGP policy, authorized-prefix filters, maximum-prefix limits, route-origin validation, BGP Roles and OTC, customer-cone validation, anomaly detection, and tested rollback address different failure modes. None is a complete retrospective cure.
  • Recovery requires external proof: A router correction or session reset is not sufficient evidence. An accountable recovery record shows withdrawals, replacement paths, residual stale routes, peer coordination, convergence, and restored reachability from independent vantage points.

The event boundary is 24 December 2004

The first discipline in an infrastructure incident review is to freeze the event. On 24 December 2004, operators observed an abnormal set of routes associated with TTNet, autonomous system 9121. The event was discussed on NANOG at the time and was later reconstructed in a NANOG presentation using public BGP data from RouteViews and RIPE collectors. [1][2] Later academic studies used the incident as a prominent example or labeled case for route-leak and routing-anomaly analysis. [3][4][5]

The public record consistently supports the central proposition: AS9121 propagated a very large portion of the routing table beyond its intended scope, and other networks accepted or redistributed enough of those announcements to cause widespread reachability problems. That is the fact pattern this article analyzes.

It is important not to inflate that proposition. Public sources do not provide a complete TTNet router configuration, every route-map, every bilateral agreement, all operator messages, or one universally authoritative timeline. They do not establish malicious intent. They do not show that every network on the internet accepted every leaked path or that every user experienced the same effect.

Numbers require particular care. The NANOG reconstruction and later literature commonly describe more than 100,000 affected prefixes and characterize the set as a majority of the global table visible at the time. [1][3][4] Those statements are meaningful, but they are measurement statements. A collector sees routes exported along the sessions feeding it. Different collectors can receive different paths, record updates at different times, suppress duplicates differently, and remain blind to routes filtered before reaching them.

The correct editorial practice is therefore attribution rather than false precision. The event was enormous by the standards of the 2004 routing table. Its exact prefix count and duration vary with vantage point and analytical method. That variability is not an excuse to minimize the incident. It is a reason to preserve the underlying BGP evidence and explain how an estimate was produced.

The date also matters. Controls standardized years later should be used as comparisons, not retroactive compliance tests. RFC 7908's route-leak taxonomy was published in 2016. [12] RFC 8212's explicit-policy default was published in 2017. [13] RFC 9234's BGP Roles and Only-to-Customer mechanism arrived in 2022. [14] These documents help identify the control problem, but they do not prove what AS9121 or its neighbors had deployed in 2004.

The defensible question is not, "Why did a 2004 network fail to comply with a 2022 standard?" It is, "Which operational invariants should have constrained what a customer or peer could announce, which organizations controlled those invariants, and what evidence would show that the route state returned to normal?"

A route leak is a relationship-policy failure

BGP distributes reachability information among autonomous systems. Each network selects routes under local policy and decides which selected routes to advertise to each neighbor. RFC 4271 defines the base protocol, message types, route attributes, selection process, and withdrawal behavior. [11] It does not encode a universal commercial or operational relationship for every session.

That relationship information matters because an internet route is not automatically eligible for every neighbor. A customer generally announces its own routes and routes for which it is authorized to provide transit. A provider can normally send broad reachability to a customer because the customer pays it for transit. A settlement-free peer generally exchanges its own and customer routes, not free transit between unrelated providers.

The terms "customer," "provider," and "peer" simplify real arrangements, but they expose the policy boundary. A route learned from one provider should not ordinarily be advertised to another provider as though the advertising network offers transit between them. A route learned from a peer should not ordinarily be exported to another peer or provider. A customer's full table should not be treated as the customer's authorized origin set.

RFC 7908 later defined a route leak as propagation beyond intended scope, usually in violation of policies associated with pairwise relationships. [12] The definition is useful here because it focuses on observable route propagation rather than presumed motive. The wrong route crossed the wrong boundary.

The TTNet event is often described as a routing-table leak because the abnormal announcements represented a vast set of routes that AS9121 should not have exported in that context. The available evidence does not require one exact internal topology to make the accountability problem clear. Whether the triggering error involved a route-map, redistribution, policy generation, session classification, or another mechanism, the externally visible failure was an export set radically inconsistent with a bounded customer or peer role.

The recipients then made local decisions. A direct neighbor accepted the routes. Some recipients may have preferred them due to path attributes, commercial preference, or path length. Some exported them further. Others may have filtered them, rejected them under a maximum-prefix limit, or selected unaffected alternatives. The event's reach was therefore an emergent property of multiple policies.

Distributed causation should not become ownerless causation. The leaker controlled what it announced. Its direct neighbors controlled what they accepted from that relationship. Each onward propagator controlled what it relayed. The relevant duties differ, but each duty is concrete.

BGP update evidence is not a universal forwarding log

Public route collectors make historical accountability possible. RouteViews archives BGP updates and routing information bases from participating peers. Its December 2004 update archive preserves data that researchers can use to reconstruct changes around the TTNet event. [9] RIPE's Routing Information Service provides another distributed measurement surface and documentation about what collectors can and cannot observe. [10]

An update record shows that a route announcement or withdrawal reached a collector over a particular session. The record can preserve prefix, AS path, origin, attributes, and timing. By comparing observations, researchers can estimate when an abnormal path appeared, how widely it propagated, and when withdrawals or replacements followed.

That is powerful evidence, but it has limits.

First, collector visibility is partial. A route may reach networks that do not feed a collector. Another route may be filtered before it reaches any public vantage point. A collector peer may export only its best path rather than every path it learned. Policy can therefore make the public record incomplete.

Second, control-plane visibility is not identical to data-plane forwarding. A router may receive an update but decline to install it. A route may enter a routing table but fail to enter the forwarding table. Traffic may follow the route but encounter congestion or a black hole later in the path. Conversely, cached sessions and path diversity can let some services remain reachable while the control plane is unstable.

Third, timestamps reflect observation. The first update in one archive is not necessarily the first bad announcement globally. The last withdrawal at one collector is not necessarily the end of stale state everywhere. Convergence is distributed.

Fourth, a prefix count depends on definitions. Analysts must decide whether to count unique prefixes, update messages, paths, origins, or route changes. They must define the incident interval and handle duplicate announcements. Different valid methods can yield different numbers.

An accountable article therefore avoids statements such as "the entire internet was down for exactly X minutes" unless a source genuinely establishes that scope. The stronger claim is also the more accurate one: AS9121 generated an extraordinary routing anomaly; the anomaly propagated widely; public measurements show the global control plane changing; and reachability was materially disrupted.

The evidence gap identifies what a complete incident package should contain. Public BGP data should be paired with the originating network's configuration diff, policy-generation logs, deployment records, alarm timeline, session states, withdrawal commands, NOC tickets, and peer communications. Direct neighbors should preserve the routes they accepted, the filters applied, any maximum-prefix events, and the reason the session remained open or was reset.

Without those private records, external investigators can reconstruct effects but cannot assign every internal action. That should be reported as an evidence boundary, not filled with speculation.

Leak is more precise than hijack

Routing incidents are often called "hijacks" in public discussion because traffic follows a path associated with the wrong network. The term can be useful for deliberate or false-origin events, but it can also imply intent that the evidence does not establish.

The TTNet record supports "route leak" as the primary term. Routes moved beyond their intended relationship scope. The event appears consistent with a serious policy or configuration failure. There is no need to allege that TTNet intended to impersonate every affected origin or deliberately intercept traffic.

The distinction matters technically. Route-origin hijacking often involves an autonomous system originating a prefix it is not authorized to originate. A path-policy leak can retain the legitimate origin but expose a route through a relationship that should not carry it. Some incidents combine elements, and public records may be ambiguous.

Route-origin validation addresses the first question: is the origin ASN authorized for this prefix under a Route Origin Authorization? RFC 6811 defines how a router can classify routes using RPKI data. [15] It does not encode every customer-provider relationship or determine whether a route with a valid origin traveled through an impermissible valley.

Relationship-aware mechanisms address a different question: given where this route was learned, should it be exported or accepted on this session? RFC 9234 formalizes BGP Roles and the Only-to-Customer attribute for a class of route-leak prevention and detection. [14] Customer-cone or ASPA-style approaches address path authorization from another angle.

Calling every event a hijack can therefore lead to incomplete remediation. An operator may create ROAs, observe that the leaked routes remain origin-valid, and conclude that the problem is solved. It is not. Origin authorization, relationship authorization, volume containment, and change safety are separate controls.

Intent still matters for legal and disciplinary findings, but intent is not required for operational containment. Filters should reject an unauthorized route set whether it was produced by a mistake, compromised automation, malicious action, or misunderstood contract. Monitoring should alert on a customer exporting most of the global table without first asking why.

This evidence-first terminology is part of the reality layer. It describes what the network asserted, accepted, and propagated. It does not turn an inference about motive into a fact.

Customer-prefix authorization should be executable

The most direct control lesson is that a customer's expected announcement set should be represented as executable policy.

If a customer is authorized to announce a defined group of prefixes, the provider can create an import filter that permits those prefixes and rejects the rest. The source of truth might include contractual records, a routing registry, RPKI origin authorizations, direct customer attestation, observed route history, and manual exceptions. Each source has limitations, but the final result must be an explicit router decision.

An allowlist is stronger than a broad denylist. A denylist tries to enumerate obviously impossible routes, such as default or reserved space, while leaving an enormous unlisted set acceptable. An allowlist starts from the routes the customer is expected to originate or transit and treats expansion as a controlled change.

The generated filter must be tested. A correct-looking registry entity can still produce incorrect configuration. An automation pipeline can merge the wrong customer, omit a prefix, accept an overly broad aggregate, or fail open when data is unavailable. Testing should compare intended resources, generated policy, and a representative set of accepted and rejected announcements.

Exceptions need ownership and expiry. A customer may need to announce a new prefix during migration, provide transit for an affiliate, or use an aggregate during an incident. An exception should identify the approving owner, evidence, affected session, prefix and path scope, start time, review time, and rollback condition. A permanent "temporary" exception is a hidden transfer of risk.

The TTNet incident shows why scale itself is an authorization signal. A network expected to announce a bounded set should not suddenly export a majority of the global table without triggering multiple controls. Even if the prefix list is stale or incomplete, the route-volume change should be extraordinary.

Customer-prefix authorization also has a reciprocal dimension. The customer should validate what it exports. It should freeze the intended route set before a change, inspect generated output, and compare actual outbound advertisements against that set. The provider should independently validate what it receives. These controls are not redundant; they reduce common-mode failure.

Written expectations do not enforce themselves. Internet Routing Registry entities, contracts, spreadsheets, and tickets are accountability records. The router implements the boundary. The practical test is whether an unauthorized route is rejected in a safe validation exercise and whether the rejection is visible to both operators.

AS-path checks and relationship policy cover different evidence

Prefix authorization asks whether the route covers an expected destination. AS-path checks ask whether the path is plausible for the relationship.

A customer may legitimately provide transit for downstream networks. In that case, a prefix-only filter tied solely to the customer's own origin ASN could reject valid service. The provider needs an authorized customer cone or another representation of which origins and paths the customer can carry.

The path representation is difficult. AS relationships change. Mergers, resellers, regional arrangements, route servers, confederations, and complex sessions resist simple classification. Public relationship inference is useful but imperfect. Private business records can be precise but may not reach routing operations in time.

That difficulty is not a reason to accept every path. It is a reason to classify confidence, constrain uncertainty, and monitor deviations. A provider can combine explicit customer declarations, registry evidence, observed stable paths, RPKI origin data, and manual review. It can reject paths containing its own ASN in an unexpected position, reserved ASNs, implausible length, or clearly unauthorized origins.

RFC 9234's BGP Roles make the session relationship explicit between two BGP speakers and define propagation rules for provider, customer, peer, route server, and route-server client roles. [14] The OTC attribute can mark routes that should subsequently travel only toward customers. The mechanism helps prevent and detect leaks that violate those relationship rules.

But BGP Roles are not a complete model of every commercial arrangement. The RFC recognizes complex relationships and warns that incorrect role configuration can itself affect propagation. The lesson is not "turn on one feature." It is "make the relationship machine-readable, confirm it with the neighbor where possible, test the result, and monitor the running route state."

For a 2004 incident, these mechanisms are retrospective comparisons. The accountability finding is older and more general: relationship policy was consequential, but too much of its enforcement depended on configuration that did not stop the abnormal export and acceptance chain.

Maximum-prefix limits provide a volume circuit breaker

A maximum-prefix limit places an upper bound on how many routes a BGP session may contribute. When the count exceeds the configured threshold, the router can warn, reject additional routes, or shut down the session depending on implementation and policy.

For an event involving more than 100,000 unexpected prefixes, a well-chosen limit is an obvious containment layer. A customer normally announcing a much smaller set should not be able to expand to near-global scale without crossing the threshold.

The control is simple in concept and delicate in operation.

Set the limit too low and legitimate growth or a deaggregation event can reset the session, causing an outage. Set it too high and it becomes decorative. Automatically restart a session without fixing the route source and it may oscillate. Warn without an owned response and the alert becomes noise.

The threshold should be based on expected route volume, growth, operational variation, and the cost of failure. It should have warning and hard-action levels. An exception should be documented and time-bounded. The response should identify whether to hold the session down, accept only the authorized subset, or coordinate a correction with the customer.

Maximum-prefix also does not prove authorization. A customer can leak a small but harmful set below the threshold. It can announce one critical more-specific route, one default, or a plausible-sized group belonging to another organization. The limit is a circuit breaker, not a substitute for prefix and path validation.

The incident demonstrates the value of independent controls. If the prefix allowlist fails, the volume limit can still contain a massive leak. If both fail, anomaly detection can compare current routes with the customer's baseline. If monitoring fails, external route alerts and peer reports can still trigger response.

An accountable operator should be able to report how many external sessions have configured warning and hard limits, how thresholds are reviewed, which sessions have exceptions, how often limits trigger, and whether exercises prove that the response protects both routing security and continuity.

Explicit import and export policy reduce ambiguous defaults

RFC 8212 changed the expected default behavior for eBGP speakers: routes should not be imported or exported when no explicit policy is configured. [13] The standard addressed a recurring class of failures in which a missing policy allowed broad propagation.

The principle applies beyond a vendor default. Every external session should have an explicit intended import and export behavior. The operator should know what happens if a route-map is absent, fails to generate, references an empty entity, or is detached during maintenance.

Fail-open behavior is attractive during availability pressure. If a registry feed is unavailable, accepting everything can keep a session up. If a policy generator errors, preserving the previous configuration may seem stale. Rejecting all routes can also break service. There is no universal response.

The accountability requirement is to choose deliberately and test the failure modes. An operator can retain the last known good policy, block only new expansions, require manual approval, or shift traffic to another path. What it should not do is discover during an incident that a missing entity silently changed "permit these routes" into "permit all."

Export policy deserves equal attention. A network may validate customer imports but accidentally advertise provider or peer routes across another relationship. It may attach the wrong community, omit a no-export control, or apply a route-map in the wrong direction. Generated outbound route sets should be compared with relationship intent before deployment.

The TTNet event is a useful test case for policy systems. Feed a representative full-table or abnormal customer route set into a lab session. Verify that the customer's own export controls reject it, that the provider's import controls reject it, that maximum-prefix contains it, and that monitoring identifies any residual route.

That test focuses on behavior. A policy document saying "customer routes are filtered" is not evidence that the current generated configuration rejects the incident class.

RPKI improves origin evidence but does not encode every leak

RPKI lets address-resource holders create cryptographically verifiable Route Origin Authorizations. A relying router can compare a BGP route's prefix and origin ASN with validated ROA data and classify the route as Valid, Invalid, or NotFound under the relevant semantics. RFC 6811 defines prefix-origin validation. [15]

This is an important improvement over unauthenticated origin claims. If a leaked announcement presents an origin ASN that the resource holder did not authorize, route-origin validation can identify and reject it under local policy.

But a path-policy leak can preserve a legitimate origin. Suppose a customer learns a valid route from one provider and leaks it to another provider while retaining the original origin ASN. The route may be origin-valid even though its propagation violates the intended relationship. RPKI origin validation does not, by itself, encode that valley.

This distinction matters for TTNet. Public summaries differ in how they describe the origin and path structure of all affected routes. An article should not claim that RPKI would have blocked every announcement without testing the actual route set against contemporaneous authorizations, which did not exist in modern form anyway.

NIST SP 800-189 recommends layered interdomain routing protections, including RPKI-based origin validation and prefix filtering. [18] Its broader framing is useful: resilient routing requires multiple mechanisms because origin security, path policy, spoofing, denial-of-service response, and operational monitoring are different problems.

RPKI also creates operational responsibilities. Operators need validated cache diversity, freshness monitoring, safe behavior during repository failure, policy for Invalid and NotFound routes, exception governance, and metrics showing actual enforcement. Creating ROAs without deploying route-origin validation leaves the evidence outside the forwarding decision.

The Heng.lu principle is visible here. Resource records are ledgers. They can establish authorization evidence and support accountability. They are not sovereign commands to routers. Running policy determines whether the evidence affects reachability.

Monitoring must compare running routes with intended routes

The TTNet leak was visible because the route state changed dramatically. A mature monitoring system should detect several dimensions of that change.

Volume monitoring asks how many prefixes a session announces, withdraws, and retains. A jump from a bounded baseline to a large fraction of the global table should be a critical event.

Origin monitoring asks whether expected prefixes changed origin ASN or whether a customer began originating address space outside its authority. RPKI and registry data can support this comparison.

Path monitoring asks whether customer, provider, and peer relationships appear in unexpected sequences. It can identify valley-like propagation, loops, sudden path shortening, or the network's own ASN in an abnormal position.

Specificity monitoring asks whether unexpected more-specific routes appeared. A small number of narrower prefixes can attract substantial traffic even when total route count remains below a limit.

Geographic and topological monitoring compares observations from multiple collectors. A route visible only in one region may be a local policy issue; a route spreading across independent upstreams indicates broader propagation.

Change correlation connects route anomalies to deployments, maintenance windows, configuration commits, and automation jobs. A global route spike seconds after a policy push should immediately identify the candidate change without requiring operators to search unrelated systems.

Monitoring should produce actionable evidence. An alert needs an owner, severity, affected session, observed route delta, expected baseline, recommended containment, and escalation path. It should preserve the route samples that triggered it.

Alerting alone is not containment. A network can detect the leak and still spend critical minutes deciding who may reset a session, which route-map to restore, how to contact peers, or whether rollback will worsen the outage. Exercises must test the operational chain.

Public monitoring provides an independent layer. Customers and critical-service operators can watch their own prefixes and origins from external vantage points. Providers can compare their view with RouteViews, RIPE RIS, or commercial feeds. Independent evidence helps detect failures in internal telemetry and confirms whether a withdrawal propagated.

Responsibility across AS9121, providers, peers, and importers

Accountability should be assigned according to control, evidence, and duty.

AS9121 had primary control over the route set it exported. Its operational review should identify the trigger, intended policy, generated configuration, approval, deployment path, detection time, containment action, withdrawal process, and remediation tests. If a customer or downstream route source initiated the table, the review should distinguish that trigger from AS9121's responsibility for accepting and exporting it.

Direct upstream providers and peers controlled the first external containment boundary. They should show the relationship assigned to the AS9121 session, expected prefix and path set, maximum-prefix policy, exceptions, alerts, and the onward export policy. If they accepted an extraordinary route volume, the review should explain which control failed, was absent, or was overridden.

Further transit networks and peers controlled later boundaries. Their responsibility depends on what they learned, from whom, and which policy was reasonable for that relationship. A downstream customer receiving a full table from its provider is in a different position from a provider accepting a full table from a small customer.

Route collector operators controlled measurement, not route propagation. Their duty is to document vantage points, timestamps, retention, and methodological limits. Researchers should make transformations reproducible where licensing and privacy allow.

Critical-service operators controlled resilience around the event. Multihoming, provider diversity, external route monitoring, out-of-band communications, and failover could reduce impact. But these operators did not control the leaked route source and should not absorb blame that belongs at interdomain policy boundaries.

End users controlled almost nothing relevant to BGP. They experienced slow connections, black holes, or service failure. Their reports can help detect impact, but they cannot validate AS paths or issue withdrawals.

This layered model avoids two errors. The first is assigning all responsibility to the originating operator while treating providers as passive pipes. The second is dispersing responsibility so broadly that no control owner remains. Each network should answer for the boundary it operated.

Recovery is a route-state transition, not a declaration

Stopping the trigger is necessary but not sufficient. BGP is distributed, and route state converges over time.

The originating network may correct a policy, withdraw routes, reset sessions, or disconnect a source. Direct neighbors must process the withdrawal or session loss, select alternatives, and export changes. Their neighbors repeat the process. Route flap damping, stale-route mechanisms, session state, timer behavior, and local policy can affect the timeline.

An operator declaration such as "the configuration was fixed" marks an internal action. It does not prove that the global route state is clean.

An accountable recovery record should answer:

  1. When did AS9121 stop exporting the abnormal route set?
  2. Which sessions were reset, filtered, or held down?
  3. When did direct neighbors observe withdrawals or replacements?
  4. How many abnormal prefixes remained at each independent collector over time?
  5. Did legitimate paths return, and were they usable in the data plane?
  6. Were stale paths or secondary leaks still visible after the primary correction?
  7. When did critical services recover from multiple regions and access networks?
  8. Which peers required direct coordination rather than automatic convergence?

The route archive can help answer some of these questions. Private session logs and route snapshots can answer more. Data-plane probes can distinguish route-table normalization from actual reachability.

Recovery evidence should also preserve uncertainty. If one collector normalized at 10:00 and another at 10:07, the report should not invent a single global recovery second. It can report a convergence interval and describe the vantage points.

The final step is a safe replay. Operators should reproduce the failure class in a lab or controlled validation environment. A representative customer should attempt to announce an oversized and unauthorized route set. The test should show which layer rejects it, which alerts fire, who responds, and how evidence is retained.

Without that test, a policy change remains a promise.

A modern control stack is layered

No single control solves every route leak, and layers should be designed to fail independently.

Explicit session policy: Import and export behavior should be explicit, with safe handling for missing or failed policy entities. RFC 8212 provides a useful default direction. [13]

Authorized-prefix filters: Direct customers should be limited to prefixes they are authorized to originate or transit, based on maintained evidence and controlled exceptions.

AS-path and customer-cone constraints: Providers should evaluate whether the path is plausible for the relationship, not only whether the origin is valid.

Maximum-prefix limits: Session-specific thresholds should warn and contain abnormal route volume before a table-scale leak propagates.

RPKI route-origin validation: Routers should use validated origin evidence under an operational policy that addresses Invalid, NotFound, cache failure, and exceptions. [15]

BGP Roles and OTC: Where supported and appropriate, mutually confirmed roles and Only-to-Customer handling can encode and enforce relationship expectations. [14]

Generated-policy testing: Automation should compare intended resources, generated configuration, and simulated announcements before deployment.

Change safety: High-risk routing changes should use staged rollout, peer review, canaries where possible, automatic stop conditions, and tested rollback.

Independent monitoring: Internal and external route views should detect volume, origin, path, and specificity anomalies and connect them to recent changes.

Incident coordination: Operators need current contacts, out-of-band channels, preauthorized containment actions, and templates for sharing affected prefixes and withdrawal evidence.

Recovery verification: Multiple control-plane and data-plane vantage points should confirm normalization.

These controls should be measured. Useful metrics include the percentage of customer sessions with generated allowlists, hard prefix limits, validated origin policy, external monitoring, current relationship classification, and recent leak-replay tests. Exception age and stale evidence are also material.

MANRS frames filtering, coordination, anti-spoofing, and global validation as operator actions. [17] NIST SP 800-189 likewise treats routing security as a set of complementary measures rather than one appliance. [18] The value of these frameworks lies in operational adoption and verification.

Registry evidence and running-code primacy

The TTNet event sits directly on the Heng.lu accountability surface: BGP routing, ASN and IP registry evidence, and peering/transit relationships.

An ASN record can identify AS9121. Address registries can identify allocations and assignees. Routing registries can publish route objects and policy. RPKI can provide signed origin authorization. These records reduce ambiguity and create an evidence trail.

They do not forward packets.

The running routers accepted and propagated a route set according to loaded configuration and live session state. If the operational policy ignored, misinterpreted, or failed to retrieve the registry evidence, the written record did not constrain the network.

This is why a registry should be understood as a ledger or recordkeeper, not a sovereign. Its legitimacy comes from accurate, unique, secure, transferable, and operationally usable records. Enforcement remains an operator responsibility.

Running-code primacy does not dismiss governance. It makes governance testable. A board can require authorized resource records, but it should also require evidence that those records produce filters and route-validation decisions. An auditor can inspect an IRR object, but it should trace that entity through the policy generator into a router and test a conflicting announcement.

The reality layer rejects two forms of advocacy copy. One says decentralization means nobody can be accountable. The other says a central registry can command the internet into correctness. Neither describes the system.

The internet is a network of independently operated systems connected by agreements and protocol sessions. Accountability comes from accurate evidence, explicit boundaries, executable policy, independent checks, and verifiable recovery.

Removing BGP export, AS9121, prefix authorization, relationship policy, collectors, and withdrawals from this article destroys the thesis. The network-control surface is not decoration. It is the subject.

What a credible remediation record would contain

A credible TTNet-class remediation package would be concrete enough for another operator or auditor to reproduce the control claim.

It would begin with an incident timeline that distinguishes the first internal trigger, first external bad route, first alert, diagnosis, containment, withdrawal, partial convergence, and verified recovery. Each timestamp would identify its source and time standard.

It would include the intended import and export policy for the relevant sessions, the actual pre-incident configuration, the change or failure that produced the leak, and the corrected configuration. Sensitive values could be redacted while preserving the logic.

It would freeze the expected customer prefix and path set, explain the references, list exceptions, and show the generated filter. It would record maximum-prefix warning and hard thresholds.

It would include representative BGP updates from internal monitors, direct peers, RouteViews, and RIPE RIS. It would explain collector limitations and counting methodology.

It would show why the original controls failed: absent filter, stale data, wrong session role, failed generation, policy ordering, broad exception, fail-open behavior, unowned alert, or another evidence-backed cause.

It would document containment decisions and peer coordination. If a session was reset, the report would show why. If routes were selectively rejected, it would show the match criteria.

It would provide a withdrawal and convergence graph by vantage point, plus data-plane checks for critical destinations.

It would report remediation ownership, due dates, completion evidence, and residual risk. A statement that "procedures were updated" would not be enough.

Finally, it would include a controlled replay showing that a full-table or abnormal customer export is rejected at multiple independent layers and that the event reaches an owned alert without escaping to production peers.

This package would not make the internet risk-free. It would make the operator's control claim falsifiable.

Governance questions should follow the route

Executives, regulators, auditors, and service buyers can ask effective questions without pretending to configure routers.

Boards should ask which changes can alter public BGP announcements, how many customer sessions lack current allowlists or hard prefix limits, how exceptions are approved, and when the last leak-replay exercise occurred.

Transit providers should ask whether relationship records are current, whether import and export policy is explicit, whether generated filters fail safely, and whether abnormal routes are contained before onward propagation.

Enterprise buyers should ask whether provider diversity is topologically independent, whether their critical prefixes are monitored externally, and whether provider incident reports include route-state evidence.

Auditors should sample live configurations and generated policies. They should trace resource evidence into route decisions and test negative cases. A screenshot of a policy portal is not proof of enforcement.

Regulators should avoid one-control mandates that confuse deployment with outcome. Requiring ROAs can improve origin evidence, but a route-leak accountability program also needs relationship policy, filtering, monitoring, coordination, and recovery tests.

Incident reviewers should not stop at "human error." That phrase does not explain why one error could export an extraordinary route set, why direct neighbors accepted it, why safeguards did not contain it, or why recovery took the time it did.

Metrics should show coverage and exceptions, not vanity. A network may report 99 percent filter coverage while leaving one high-volume customer unrestricted. Risk follows the exposed boundary, not the average.

Governance should also protect truthful uncertainty. Operators should be expected to disclose which parts of the route state they could not reconstruct, which data was not retained, and which counterfactual controls were not tested.

The accountability test is bounded propagation and verified withdrawal

The 24 December 2004 TTNet event remains relevant because it exposed a control problem that still exists wherever BGP relationships are implicit, filters are broad, and route evidence is disconnected from running policy.

AS9121 propagated an extraordinary route set. Other autonomous systems accepted and redistributed enough of it to turn a local policy failure into a distributed reachability event. Public route collectors preserved part of the control-plane record. Later analysis made the event useful as a benchmark for leak detection.

The public evidence does not justify malicious-intent claims, one exact universal prefix count, or a complete reconstruction of every operator decision. Those limitations should remain visible.

Responsibility was layered. The exporting network controlled its route generation and announcements. Direct providers and peers controlled the first external acceptance boundary. Further networks controlled onward propagation. Monitoring operators controlled evidence quality. Service operators controlled resilience. End users bore impact without control over the route state.

Modern remediation is also layered. Authorized-prefix filters, AS-path constraints, maximum-prefix limits, explicit policy, RPKI validation, BGP Roles, policy testing, independent monitoring, coordinated response, and convergence verification address different parts of the failure chain.

The Heng.lu doctrine supplies the final distinction. Registry and routing-policy records can establish who is expected to originate resources and what relationships are intended. They are indispensable accountability ledgers. The running routers still determine reachability.

A credible operator should therefore be able to prove three things.

First, it knows what a customer or peer is authorized to announce.

Second, its running systems reject or contain announcements that exceed that scope, even when one policy source or automation component fails.

Third, when bad state escapes, it can withdraw it quickly and show from independent vantage points that the internet returned to an authorized route state.

That is the customer-prefix filtering accountability test made visible by the 2004 TTNet leak. The standard is not perfect prevention. It is bounded propagation, assigned control, retained evidence, and verified recovery.

Sources

  1. NANOG 34, "Route Leakage and Misorigination: TTNet (AS9121), December 2004"
  2. NANOG mailing-list archive, December 2004 incident discussion
  3. Electronics 2024, route-leak detection study using the TTNet incident
  4. IEEE Transactions on Network and Service Management, BGP anomaly analysis
  5. University of Cambridge Technical Report 898, route-leak analysis
  6. KTH routing-security lecture, AS9121 historical event
  7. bgp.tools, current AS9121 public routing view
  8. RIPEstat, current AS9121 public routing evidence
  9. RouteViews, December 2004 BGP update archive
  10. RIPE NCC, Routing Information Service documentation
  11. RFC 4271, A Border Gateway Protocol 4
  12. RFC 7908, Problem Definition and Classification of BGP Route Leaks
  13. RFC 8212, Default External BGP Route Propagation Behavior Without Policies
  14. RFC 9234, Route Leak Prevention and Detection Using Roles in UPDATE and OPEN Messages
  15. RFC 6811, BGP Prefix Origin Validation
  16. RFC 7454, BGP Operations and Security
  17. MANRS, Network Operators Actions
  18. NIST SP 800-189, Resilient Interdomain Traffic Exchange