Summary
- A valid optional transitive BGP attribute became corrupted when affected IOS XR systems propagated it, causing downstream session resets and wider routing instability.
- Accountability follows separate control points: experiment scope and monitoring, product behavior and repair, operator continuity controls, and protocol error handling.
On 27 August 2010, RIPE NCC staff operating the Routing Information Service, or RIS, worked with a Duke University research group on a live Border Gateway Protocol experiment. The researchers were studying a secure-routing design that carried certification information in an optional transitive path attribute. At 08:41 UTC, RIPE NCC originated 93.175.144.0/24 from RIS AS12654 through connections at AMS-IX and GN-IX. The route was withdrawn as planned at 09:08 UTC.[1]
The announcement was unusual because the attribute was new on the public Internet. Unusual, however, did not mean malformed. RIPE NCC and Cisco described the initiating attribute as valid or standards-compliant.[1][2] That distinction determines how the event should be analyzed. The experiment did not simply demonstrate that routers reject invalid input. It showed that a standards-compliant input could pass formal checks, encounter a defect in deployed software, become corrupted during propagation and then activate severe error handling elsewhere.
Cisco reported that affected IOS XR systems mishandled the valid but unrecognized transitive attribute while sending it onward. A neighboring router could receive the resulting corrupted UPDATE and reset the BGP peering session. Because a BGP session carries many routes, a response to one defective UPDATE could temporarily remove unrelated valid reachability information. If the router responsible for outbound corruption did not detect its own error, it could advertise the problematic information again after the session returned, renewing the instability.[2][3]
RIPE NCC’s measurements support a bounded but material impact finding. The event produced update rates as high as twenty times the surrounding baseline. RIPE estimated that an additional 0.5 percent of prefixes became completely unreachable for longer than normal. The proportion of unstable prefixes peaked at 1.4 percent, representing nearly 4,500 prefixes and approximately nine times the ordinary level observed around the event. Results varied by collector and location, with especially intense update activity visible at the Vienna collector.[1]
Those figures are significant, but they are not a license to claim that a known percentage of users, routers, networks or traffic disappeared. Control-plane observations count routing behavior, not people. RIPE’s DNSMON analysis found no root-server failure. It did find limited query loss for some monitored domains and more visible problems at portions of the .si and .fr authoritative infrastructure, while redundant servers continued answering.[1] Contemporary operator comments mentioned access problems and exchange-traffic changes, but such reports do not form a complete quantitative account.[4][6]
The accountability lesson is therefore narrower and more useful than a story of one actor “breaking the Internet.” RIPE NCC controlled whether its measurement infrastructure originated an Internet-visible experimental route, along with timing, notice, monitoring and withdrawal. Duke researchers controlled the research design and research-side implementation. Cisco controlled IOS XR’s handling of an unrecognized valid attribute, its product testing, disclosure and maintenance fixes. Network operators controlled their installed software, routing policies, peer protections, monitoring and restoration within their own networks.
The protocol rules of the period supplied an amplification mechanism by allowing one malformed UPDATE to cause session-level failure.
The repairs also belonged to separate layers. Cisco issued an advisory on the day of the event and prepared maintenance upgrades.[2] RIPE NCC preserved evidence, supplied information to the vendor, published an analysis and committed to stricter controls for future cooperative experiments.[1] Its Executive Board later supported continued research while emphasizing appropriate communication.[5] Product correction reduced the implementation risk; stronger experiment governance reduced the risk of discovering an unknown interaction through an unnecessarily broad public failure domain.
Later standards help explain how the industry learned from this class of failure, but they cannot be used as proof of what was deployed or required in August 2010. RFC 7606 later promoted narrower handling of malformed BGP UPDATEs because resetting an entire session can discard many valid routes.[13] Operational recommendations concerning explicit external policy, route-leak prevention, origin validation, RPKI and BGPsec illuminate adjacent controls, but none retroactively establishes negligence, and none alone corrects the precise outbound-corruption defect exposed by this event.[14][15][16][17][18][19][20]
The central conclusion is restrained: a valid experimental announcement triggered a chain in which an implementation defect corrupted information, session-level error handling amplified that corruption and insufficiently bounded live-experiment controls exposed the interaction on the public Internet. Accountability follows the allocation of control, not the simplicity of the first visible trigger.
1. Why this event deserves a narrow analysis
BGP incidents are often compressed into labels such as leak, hijack, outage or misconfiguration. Those labels can be useful when the evidence fits them, but they can also erase the mechanism that actually matters. This event requires a more precise boundary.
The subject is only the RIPE NCC–Duke University experiment of 27 August 2010. It is not an account of later route leaks, later hijacks or every subsequent BGP failure. RFC 7908’s later route-leak taxonomy is useful for distinguishing categories, but it does not justify relabelling this experiment as a conventional route leak without evidence that the specified leak conditions occurred.[16]
The experiment originated a deliberately constructed route carrying a new optional transitive attribute. Its public visibility was intentional. The resulting instability was not. That combination creates three questions that should not be collapsed into one:
- Was the initiating BGP UPDATE valid?
- Which component changed valid information into corrupted information?
- Which controls allowed the resulting failure to spread beyond a narrowly bounded test?
The available record answers the first two questions with relatively high confidence. RIPE NCC and Cisco both characterized the original attribute as valid or standards-compliant, while Cisco identified an IOS XR vulnerability involving corruption during propagation.[1][2] NVD records the product issue as CVE-2010-3035.[3]
The third question is distributed. The researchers did not control each deployed router. Cisco did not decide that RIS would originate the experimental route. Individual operators did not design the experiment or the affected implementation. Internet exchange connections did not, by their mere presence, prove approval of every attribute carried across them. Attribution must therefore follow the actual control points.
This event matters because it breaks an easy but unsafe assumption: that standards compliance at the origin is enough to establish operational safety across a heterogeneous routing system. Compliance is necessary, but deployed behavior decides whether a packet or UPDATE survives contact with running code. The experiment passed one kind of validity boundary and failed another.
RIS itself exists to collect and expose routing information from multiple observation points.[7][8] RIPEstat and RIPE Database records provide identifying context for AS12654, but an autonomous-system record cannot reveal every software behavior encountered along a propagated path.[9][10] Route-collector systems such as RIS and Route Views are valuable precisely because no single registry record, peer statement or local log can describe the entire interdomain routing system.[11]
The lesson is not that standards are irrelevant. It is that a standard describes required behavior, while accountability requires evidence that implementations, deployments and operational controls produce that behavior under real conditions.
2. Chronology: from a scheduled announcement to an unintended disturbance
Before 08:41 UTC: design and pre-announcement checking
The Duke research group was studying a secure-routing design in which certification information would travel in an optional transitive BGP path attribute. Duke supplied a modified Quagga implementation within the research effort. The pre-announcement checks established that the attribute had an acceptable protocol form, and a second Quagga instance did not reproduce the behavior later seen in affected deployed equipment.[1]
That result was informative but incomplete. It showed that the initiating implementation and a similar receiving environment could process the attribute. It did not show that every router family, software release or forwarding path on the public Internet would preserve the attribute correctly.
This was the critical detection gap. The test path assessed formal structure and limited implementation behavior. It did not reproduce the heterogeneous installed base through which a transitive attribute could travel. Most importantly, it did not expose a defect in which a router accepted a valid unknown attribute but corrupted it while sending it onward.
The public record does not establish the full approval record, every contemplated risk or every control discussed before the event. It would therefore be inappropriate to invent an undocumented decision process. What can be said is that the controls used before the announcement did not detect the interaction that produced the public disturbance.
08:41 UTC: the route becomes visible
At 08:41 UTC on 27 August 2010, RIS AS12654 began announcing 93.175.144.0/24 with the experimental optional transitive attribute. The announcement travelled through RIPE NCC’s connections at AMS-IX and GN-IX.[1]
This moment was the triggering event. “Trigger,” however, is not synonymous with “root cause.” A trigger is the event that activates a latent condition. If a valid input encounters a defective implementation, the valid input starts the observed sequence, but the defect explains why the sequence departs from the specified behavior.
The route’s public visibility also created the experiment-governance exposure. A test on a closed or tightly bounded system can reveal a defect without allowing it to traverse a broad set of autonomous networks. Once the route entered ordinary interdomain propagation, the outcome depended on software and policies outside the initiating parties’ direct control.
During propagation: valid information is corrupted
Affected IOS XR systems received an attribute they did not recognize. Under BGP’s optional-transitive design, lack of recognition was not itself grounds to reject the attribute. The relevant behavior was to preserve and propagate it.
Cisco’s account identified a failure in that propagation path: affected systems corrupted the otherwise valid attribute while sending it to a neighbor.[2] The available evidence does not justify inventing the precise byte-level mutation for every affected path. The reliable finding is functional: valid unfamiliar information entered an affected implementation, and corrupted information emerged on propagation.
That distinction locates the product defect more accurately than saying the router merely “did not support” the experimental feature. Optional transitivity exists so that a router can carry an attribute without understanding its full semantics. A conforming implementation does not need to act on the certification information. It does need to preserve the attribute correctly if it propagates it.
The corruption then crossed an implementation boundary. A downstream neighbor received an UPDATE that was no longer equivalent to the valid one originated by the experiment.
Downstream reception: one UPDATE threatens a whole session
Under the error-handling rules associated with the standards baseline of the period, a malformed path attribute could produce an UPDATE Message Error and closure of the BGP session.[12] This response was severe by design: a router that could not safely interpret routing information protected itself by terminating the session.
The operational side effect was broad. A BGP session ordinarily carries many routes, not only the experimental prefix. Closing the session could therefore withdraw unrelated valid routes learned from that peer. Networks would then search for alternatives, exchange new UPDATEs and reconverge.
A further repetition mechanism was possible. The router that corrupted the outbound attribute did not necessarily identify itself as the source of the corruption. After a neighboring session recovered, the same route could be advertised again. The same corruption could recur, the downstream system could reset again and routing instability could renew.[2]
The failure chain was consequently larger than the experimental prefix:
- One valid but unfamiliar attribute entered an affected router.
- The router corrupted it during onward propagation.
- A neighbor received a malformed UPDATE.
- The neighbor could terminate the BGP session.
- Routes unrelated to the experiment could disappear from that adjacency.
- Reconvergence produced a larger volume of updates.
- Re-advertisement could repeat the sequence.
This is why the event became an accountability test rather than merely an interoperability curiosity. Each stage was controlled by a different component or organization.
09:08 UTC: scheduled withdrawal
RIPE NCC withdrew the experimental announcement at 09:08 UTC, as planned.[1] The route had therefore been originated for approximately twenty-seven minutes.
Withdrawal was necessary, but a withdrawal is not an instantaneous eraser. BGP is distributed. Updates already accepted or propagated must travel through other sessions, while routers recalculate paths and restore adjacencies. A planned withdrawal can stop continued origination at the initiating point without immediately canceling every copy, queued update, reset or reconvergence process already under way.
RIPE described the unintended operational impact as lasting about thirty minutes. Most instability returned toward normal roughly twenty minutes after the experiment, rather than ending at the exact second of withdrawal.[1] Those descriptions should be treated as measured operational intervals, not converted into a claim that every affected path recovered simultaneously.
After withdrawal: operator response and evidence collection
Operators observed and responded from their own networks. Contemporary mailing-list discussion included reports of access problems, routing changes and exchange-traffic effects.[4][6] Such comments are useful evidence that the disturbance was operationally visible, but they have strict limits. They do not enumerate every affected autonomous system, normalize observations across locations or supply a complete count of lost traffic.
RIPE NCC retained experiment data and provided collected evidence to Cisco. That preservation mattered because the event crossed organizational boundaries. An origin-side log alone could show what RIS sent, but not necessarily what an intermediate implementation emitted. A vendor report alone could explain a defect, but not quantify Internet-wide observations. A credible reconstruction required evidence from more than one control domain.
22:00 UTC: Cisco’s advisory
Cisco published its IOS XR advisory at 22:00 UTC on 27 August 2010.[2] The issue was recorded as CVE-2010-3035, and Cisco prepared software maintenance upgrades.[3]
The same-day advisory established a public product-level response. It did not, by itself, prove which release was installed at every affected operator, how many devices encountered the route or whether every operator had an immediately usable mitigation. Those remain unknown in the bounded record.
31 August and afterward: public analysis and governance response
RIPE NCC published its incident and measurement analysis on 31 August 2010.[1] The account described the experiment, the implementation interaction, observed routing effects, DNSMON findings and prospective changes to future cooperative experiments.
RIPE NCC said future experiments would receive stricter treatment, including comprehensive impact assessment, sufficient advance notice to operators and responsible handling of vulnerabilities.[1] Its Executive Board later supported continued experimentation while emphasizing suitable communication.[5]
That response did not deny the value of research. It recognized that a useful research objective does not eliminate the need to bound operational exposure. The governance repair was therefore not “stop experimenting.” It was to make the authority to initiate a public experiment conditional on clearer risk, notice, containment and response controls.
3. What an optional transitive attribute is—and why “unknown” did not mean “invalid”
BGP path attributes carry information associated with a route. Some are well-known and expected to be understood by BGP implementations. Others are optional. A separate distinction concerns transitivity: whether an attribute should continue across BGP speakers even when an intermediate implementation does not recognize its meaning.
RFC 4271 defines the relevant behavior. An unrecognized optional non-transitive attribute does not need to be forwarded. An unrecognized optional transitive attribute is different. It is accepted and passed to other BGP peers, with the Partial bit used to indicate that an intermediate system did not fully recognize the attribute.[12]
That mechanism supports extension. Without it, every autonomous system along a path would need simultaneous software support before a new transitive feature could cross the Internet. Optional transitivity permits incremental deployment: a router may transport information without interpreting it.
The design creates a strict implementation responsibility. A router that does not understand an optional transitive attribute must still handle its representation safely. In practical terms, it must not turn valid opaque information into malformed information.
Four concepts must remain separate:
Recognition. Does the router understand the attribute’s semantics?
Acceptance. Does the attribute have a form that the router can safely receive under the protocol rules?
Propagation. Must or may the router pass the attribute onward?
Mutation. Does the router alter the attribute, and if so, is that alteration permitted and correctly encoded?
The 2010 event did not require affected IOS XR systems to understand the research certification scheme. The defect concerned propagation. Cisco’s account was that affected systems corrupted the valid unrecognized transitive attribute while sending it onward.[2]
This explains why a second Quagga instance was not enough to predict the incident. Two implementations can agree about the form of an attribute while a third contains a defect in a different code path. Receiving, storing and serializing an unknown attribute may involve separate operations. Passing a conformance check at entry does not prove correct behavior at exit.
The event also demonstrates the difference between syntactic validity and end-to-end operational safety. The initiating UPDATE could be valid at origin. An intermediate system could then produce an invalid representation. A downstream system could respond correctly according to the rules available to it and still create a damaging operational consequence by closing an entire session.
No single layer alone explains the impact:
- The experiment supplied the unfamiliar input.
- The affected product supplied the corruption.
- The downstream error response supplied session loss.
- BGP reconvergence supplied update amplification.
- Public propagation supplied the failure domain.
Calling the original attribute malformed would erase the product defect unless packet evidence proves otherwise. Calling the entire event only a product bug would erase the decision to expose an uncertain interaction on the live Internet. Calling it only a harsh protocol response would erase the implementation that generated the malformed downstream UPDATE.
The accurate description is a chained failure with distinct control owners.
4. From one corrupted UPDATE to wider routing instability
BGP distributes reachability between autonomous systems. When a peering session closes, routes learned exclusively or preferentially through that session can be withdrawn from the local routing table. The router may select alternatives and announce those changes to other peers. Those peers then repeat their own selection processes.
This means that an error attached to one route can create changes involving many routes if the response removes a whole adjacency. The experimental prefix did not need to be the destination sought by affected users. A reset could disturb other reachability learned across the same session.
The event’s update volume is consistent with this amplification mechanism. RIPE observed update rates as high as twenty times the surrounding baseline.[1] That is a control-plane measurement: routers were exchanging substantially more routing changes. It does not directly state how much application traffic was lost, but it does show that the disturbance extended beyond one quiet rejection of one prefix.
Session restoration could also create recurrence. If the upstream affected router retained the route and repeated its defective propagation when the session returned, the neighbor could encounter the malformed UPDATE again. The resulting cycle would combine:
- session establishment;
- route advertisement;
- corruption during propagation;
- malformed UPDATE reception;
- session closure;
- route withdrawal and reconvergence; and
- renewed establishment.
Not every peer or path necessarily experienced every step. The evidence supports a mechanism that could repeat and observations of elevated instability; it does not supply a complete packet-by-packet record for every autonomous system.
This distinction matters when assigning impact. A route collector sees advertisements and withdrawals at its observation point. It does not see every forwarding decision, every user session or every dropped packet. Different collectors see different slices of the routing system. The especially intense activity at the Vienna collector illustrates that the effect was uneven.[1]
Unevenness is not a flaw in the measurement. It is a property of the Internet’s topology and policy. Autonomous systems choose routes locally. They have different peers, software, filters and alternatives. A defective advertisement can pass through one path, be blocked on another and never be selected on a third.
The correct analytical move is therefore not to extrapolate one collector’s peak to the whole Internet. It is to combine collectors, describe the distribution and preserve the limits of inference.
5. Bounded impact: what the evidence supports
RIPE’s measurements provide three principal indicators.
First, routing-update rates reached as much as twenty times the surrounding baseline.[1] This demonstrates exceptional control-plane activity during the event window. The phrase “as high as” matters: it describes a peak, not a uniform rate at every collector or for the entire period.
Second, RIPE estimated that an additional 0.5 percent of prefixes became completely unreachable for longer than normal.[1] This is a prefix-level visibility measure. It should not be translated into 0.5 percent of users, traffic, routers or economic activity. Prefixes vary greatly in size, use and traffic, and route collectors do not observe every forwarding path.
Third, the share of unstable prefixes peaked at 1.4 percent. RIPE associated that peak with nearly 4,500 prefixes, about nine times the usual level.[1] “Unstable” is not identical to “universally unreachable.” A prefix may experience repeated path changes while remaining reachable from some places.
These findings support the conclusion that the event caused material, measurable and distributed routing instability. They do not support an assertion that 1.4 percent of the Internet went entirely offline.
A contemporary shorthand that the event affected roughly one percent of the Internet may capture the order of magnitude of some measurements, but it is less precise than RIPE’s separate indicators. It should not replace them. The evidence distinguishes additional complete unreachability, observed instability and update volume.
Geographic and collector variation
The effects varied by location and collector. The Vienna collector showed especially high update activity.[1] Variation can reflect topology, peer selection, exposure to affected implementations and the availability of alternate paths.
A collector’s location is not a direct map of user impact in that city or country. BGP observation points receive routes from participating peers. Their view may include paths serving remote networks, and local users may follow paths not visible to the collector. The collector evidence is strong for routing behavior and weaker for assigning a geographic count of affected people.
DNS observations
RIPE used DNSMON to examine whether the routing disturbance produced visible DNS effects. It did not find a failure of the root-server system.[1] That negative finding is important because broad routing instability does not automatically imply failure of every critical service.
The analysis did observe limited query loss for some monitored domains and more noticeable difficulty at portions of the .si and .fr authoritative infrastructure. Redundant servers continued to answer.[1] The evidence therefore supports partial and uneven DNS effects, not universal DNS failure.
The continued availability of redundant servers is also a reminder that routing accountability includes service architecture. A routing disturbance can reach one authoritative server path while another remains reachable. Redundancy does not eliminate the routing defect, but it can prevent a component failure from becoming complete service failure.
Operator reports
Contemporary operational forums recorded reports of access disruption, routing reactions and traffic changes.[4][6] These reports help establish that the incident was visible outside the initiating institutions. They can also identify questions for further investigation.
They are not a substitute for normalized measurement. A traffic drop at one exchange or operator may reflect rerouting, loss, precautionary policy changes or another local response. Without matched baselines, topology and traffic records, it cannot be converted into a total Internet impact figure.
Claims the evidence does not support
The bounded record does not establish:
- a complete list of affected IOS XR releases as deployed at the time;
- an exact number of affected routers or devices;
- every autonomous system that reset a session;
- an exact number of affected users;
- total application traffic lost;
- universal failure of the experimental prefix;
- failure of the DNS root;
- malicious intent by RIPE NCC, Duke, Cisco or operators;
- a complete account of every pre-event approval;
- that every operator had an available mitigation before the event; or
- legal liability.
These are not minor disclaimers. They define the difference between evidence-based infrastructure reporting and an outage story built from unsupported multiplication.
6. Accountability through control allocation
Accountability is strongest when it asks who controlled each consequential decision, implementation and recovery action. It becomes weaker when it treats proximity to the first visible event as proof of sole responsibility.
| Control domain | What the entity controlled | What the entity did not control | Evidence needed for a stronger assessment |
|---|---|---|---|
| RIPE NCC | Use of RIS infrastructure, Internet-visible origination, timing, communication, monitoring, withdrawal, evidence retention and future experiment policy | Software behavior on every external router and every operator’s recovery | Approval record, risk assessment, notice plan, monitoring thresholds, withdrawal criteria and retained observations |
| Duke researchers | Research design, experimental attribute construction, research-side Quagga changes and research-side testing | Deployed IOS XR code, downstream session policies and operator software deployment | Test vectors, generated UPDATE bytes, research implementation records and the scope of interoperability testing |
| Cisco | IOS XR parsing, storage and propagation behavior; product test coverage; disclosure; maintenance fixes | The decision to originate the experiment and operators’ installation schedules | Defect analysis, affected-release matrix, regression results, fixed-code evidence and deployment guidance |
| Network operators | Installed software, maintenance, import and export policy, peer controls, filtering, monitoring and restoration inside their networks | Experimental design, upstream vendor code and the complete global propagation path | Device logs, packet captures, configurations, software versions, session history and restoration records |
| Downstream BGP implementations | Local handling of the malformed UPDATE under the rules they implemented | The creation of the original valid attribute or upstream corruption | UPDATE error logs, session notifications and proof of narrower handling where supported |
| Internet exchanges | Connectivity through which participating networks exchanged routes | By default, the content and correctness of each entity’s BGP announcement | Evidence of any specific route-server, filtering or operational role before assigning more control |
RIPE NCC’s control
RIPE NCC controlled the act that introduced the experimental route into public propagation. RIS AS12654 was the origin used for the test, and the announcement traversed RIPE NCC connections at AMS-IX and GN-IX.[1] RIPE NCC also controlled the planned withdrawal, its collection of evidence and its future rules for similar cooperative research.
That control establishes experiment-governance accountability. It does not establish that RIPE NCC created the product defect. The initiating attribute was described as valid. The relevant question for RIPE NCC is not whether it should have predicted the exact undocumented failure with certainty. It is whether the experiment’s uncertainty was assessed, communicated, monitored and contained in proportion to its possible public reach.
The later commitment to more comprehensive impact assessment, advance operator notice and responsible vulnerability handling indicates that RIPE NCC itself identified governance improvements.[1] The Executive Board’s support for continued experimentation with appropriate communication reinforces a distinction between the legitimacy of research and the adequacy of its operational controls.[5]
Duke’s control
Duke researchers controlled the secure-routing research design and supplied the Quagga modification used within their scope. Their work helped create the valid experimental input. The available record does not show that they controlled IOS XR’s internal handling or downstream routers’ error response.
Research-side accountability concerns the design assumptions and breadth of interoperability testing. A second Quagga instance could demonstrate behavior in a similar software environment. It could not establish safety across all relevant deployed implementations.
The public evidence does not reveal the complete division of pre-event decisions between Duke and RIPE NCC. It would be improper to invent one. Any more granular assignment would require experiment plans, test records and communications that identify who approved the public propagation conditions.
Cisco’s control
Cisco controlled the affected IOS XR implementation. Its advisory identified corruption of a valid unrecognized transitive attribute during propagation.[2] That behavior sits within the product control domain: parsing, retention, serialization, attribute-flag treatment and regression testing.
Cisco also controlled its disclosure and maintenance response. The advisory appeared on the day of the incident, and software maintenance upgrades were prepared.[2] CVE-2010-3035 provides the public vulnerability identifier.[3]
Product accountability should still remain evidence-based. The record does not establish every deployed release, the number of affected devices or whether the defect had been discovered earlier. A stronger assessment would require release-specific testing, defect history and installation evidence.
Operators’ control
Each network operator controlled a local portion of the system: software selection and installation, maintenance timing, peering policy, filters, monitoring, session protection and restoration. These controls could affect exposure and recovery.
That does not make operators responsible for predicting an unknown vendor corruption defect. Nor does it show that every operator had an available patch or configuration mitigation before the experiment. Operator accountability is conditional on what was knowable and controllable at the relevant time.
After disclosure, the evidence required for continuing assurance changes. Operators can be asked to identify affected releases, apply fixes, test behavior and retain proof. Before disclosure, claims about unreasonable inaction would require evidence that a risk and feasible mitigation were already known.
Why exchanges should not be assigned an invented role
The experiment used connections at AMS-IX and GN-IX.[1] That fact establishes a propagation path. It does not, without additional evidence, establish that either exchange designed the experiment, approved the attribute, operated an affected router or controlled entity export policies.
Infrastructure reporting often confuses physical or logical transit with decision authority. A named exchange can be part of the route’s path without being the actor that originated, corrupted or accepted the UPDATE. Accountability should not be inferred from topology alone.
7. Product repair and experiment-governance repair are different
A complete response required two repair tracks.
Product repair
The product defect was the corruption of a valid unrecognized transitive attribute during propagation by affected IOS XR systems. The direct repair belonged in the software and its tests.
A credible product repair would demonstrate that:
- a valid unknown optional transitive attribute can be received;
- it is stored without destructive mutation;
- it is propagated in the form required by the protocol;
- the relevant attribute flags and length fields remain consistent;
- repeated session establishment does not recreate corruption;
- malformed variants are contained according to the supported error-handling behavior;
- regression tests cover both recognition and opaque propagation paths; and
- the corrected release is identifiable to operators.
Cisco’s advisory and maintenance upgrades were the immediate public actions addressing this layer.[2] An advisory communicates the defect; an upgrade changes the implementation. The two are related but not interchangeable.
Verification also requires deployment evidence. A vendor can prove that a corrected build passes regression tests, while an operator can prove which build is running on a particular router. Neither record alone establishes both product correction and field adoption.
Experiment-governance repair
The governance defect was not that research occurred. It was that an uncertain interaction was tested through public routing infrastructure without controls sufficient to prevent or rapidly limit the observed blast radius.
RIPE NCC’s response identified stricter future requirements: comprehensive impact assessment, enough advance notice for operators and responsible treatment of vulnerabilities.[1] These address decisions made before and during an experiment.
A credible governance repair would include:
- a clearly bounded technical objective;
- identification of every attribute and route to be originated;
- a documented propagation boundary or explanation of why broader propagation is required;
- heterogeneous implementation testing appropriate to the risk;
- advance communication to affected operators where feasible;
- a defined test window;
- real-time route-collector observation;
- data-plane or service probes where relevant;
- quantitative kill criteria;
- an authorized person able to withdraw immediately;
- a rehearsed withdrawal procedure;
- criteria for contacting vendors;
- preserved pre-event and post-event data; and
- a public incident account when unintended external impact occurs.
Governance controls cannot guarantee that an unknown defect will never appear. Their purpose is to reduce the probability that discovery produces uncontrolled external consequences and to shorten the time between detection and containment.
Why one repair cannot substitute for the other
If Cisco corrected IOS XR but experiment controls remained unchanged, a later experiment could expose a different unknown defect in another implementation. The specific product risk would decline, while the discovery risk remained.
If RIPE NCC strengthened experiment controls but affected software remained uncorrected, ordinary Internet traffic carrying another valid unfamiliar transitive attribute could still encounter the latent defect. The public-test risk would decline, while the product risk remained.
The event therefore requires two independent closure questions:
- Is the implementation defect corrected and deployed where relevant?
- Are future live experiments bounded, observable and governed in a way that matches their uncertainty?
A report that answers only one has not demonstrated complete repair.
8. Later standards as analytical context—not retroactive judgment
Standards published after August 2010 help describe better containment and policy practices. They do not prove that those practices were deployed during the event, and they cannot retroactively convert later recommendations into a finding of negligence.
RFC 4271: the historical baseline
RFC 4271 describes BGP-4, including optional transitive attributes and error handling.[12] Its propagation rules explain why an unrecognized transitive attribute should be carried onward. Its UPDATE error handling also helps explain why a malformed attribute could lead to session termination.
That combination produced a dangerous interaction. Extensibility depended on safe opaque propagation, while malformed input could activate a broad response. When an intermediate implementation corrupted opaque information, the downstream system faced an error condition with consequences larger than the one route.
RFC 7606: narrowing the failure domain
RFC 7606 later revised BGP UPDATE error handling because session reset can discard large numbers of valid routes and cause substantial routing disruption.[13] It generally promotes narrower responses, including treating affected routes as withdrawn in defined cases, rather than automatically destroying the entire session.
Applied as analytical context, this shows how the failure domain can be reduced. If a malformed advertisement can be contained to the affected route while the session and unrelated routes remain, one corrupted attribute has less power to destabilize an adjacency.
It would be inaccurate to say RFC 7606 was the rule governing the 2010 event. It was published later. It would also be inaccurate to assume every current implementation applies every recommendation uniformly. The RFC explains an architectural repair direction; deployment evidence is still required.
RFC 7454 and RFC 8212: explicit external policy
RFC 7454 collects operational security recommendations for BGP, while RFC 8212 establishes an explicit-policy expectation for external BGP announcements and acceptance.[14][15] Together they reinforce a basic control principle: external routes should not be exchanged merely because a session exists.
Explicit import and export policies can reduce accidental propagation and make intended relationships auditable. In a bounded experiment, carefully scoped policy could help limit which peers receive a test route.
These measures do not directly correct a router that corrupts an attribute it is required to propagate. They operate at the policy boundary, not inside the faulty serialization path. They may reduce exposure, but only if the route or session can be distinguished and constrained without defeating the experiment’s legitimate objective.
RFC 7908: route-leak taxonomy
RFC 7908 describes types of route leaks.[16] It is useful here mainly as a boundary against loose terminology. The 2010 event involved an intentionally originated experimental route and an implementation defect affecting an optional transitive attribute. The available evidence should not be stretched to place it into a later leak category without matching the category’s conditions.
Taxonomy supports accountability when it prevents unrelated mechanisms from being merged. It undermines accountability when a familiar label replaces causal analysis.
RPKI origin validation
RFC 6480 describes the Resource Public Key Infrastructure architecture, and RFC 6811 defines BGP prefix-origin validation.[17][18] Origin validation asks whether an origin autonomous system is authorized by a relevant Route Origin Authorization for a prefix.
That control addresses a different question from the one exposed here. A route can have an acceptable origin relationship while carrying an attribute that an intermediate product later corrupts. Origin validation does not prove that every path attribute is correctly encoded or preserved.
No conclusion about the experiment’s actual RPKI state is necessary. The analytical point is limited: origin validation, by itself, would not test the affected outbound handling path.
BGPsec
RFC 8205 specifies BGPsec path validation.[19] BGPsec addresses cryptographic protection of path information in a defined architecture. It is related to the broader goal of secure routing studied by the Duke group, but it is not evidence about what was deployed during this 2010 experiment.
Nor should it be presented as an automatic correction for every implementation defect. Security mechanisms are themselves implemented in software. Safe parsing, serialization, failure containment and interoperability testing remain necessary.
NIST routing-security guidance
NIST SP 800-189 provides later guidance for securing interdomain traffic exchange, including routing protections and operational practices.[20] It is useful for structuring present-day expectations around filtering, monitoring, validation and response.
It does not establish a 2010 legal duty or prove what any entity knew at the time. Its proper use is prospective: to ask what evidence a network should now retain and which controls can reduce similar failure chains.
9. Counterfactuals: which changed fact would have reduced the impact?
Counterfactual analysis is useful only when each scenario changes a defined condition and preserves the rest of the evidence. It cannot prove what would certainly have happened, but it can identify high-value controls.
Counterfactual 1: IOS XR preserves the attribute correctly
Change one fact: affected IOS XR systems receive the valid unrecognized transitive attribute and propagate it without corruption.
The downstream malformed UPDATE does not arise from that product path. The session-reset mechanism attributed to the corrupted UPDATE is therefore not activated by this defect. The announcement remains unusual and experimental, but the documented failure chain is interrupted at its principal implementation point.
This is the strongest product counterfactual because it removes the identified corruption mechanism. It does not prove that no other implementation would have reacted badly.
Counterfactual 2: heterogeneous testing reproduces the deployed behavior
Change one fact: pre-public testing includes a sufficiently representative affected implementation and triggers the outbound corruption.
The defect can be investigated before the route enters broad public propagation. Cisco can receive the test case, while RIPE NCC and Duke can decide whether to postpone, constrain or redesign the experiment.
The limitation is representativeness. No laboratory can reproduce every Internet path. The value lies in expanding beyond two similar Quagga endpoints and specifically testing opaque receipt-and-propagation behavior across distinct implementations.
Counterfactual 3: the experiment is conducted inside a bounded routing environment
Change one fact: the same attribute and affected software interact in a closed or tightly controlled test environment rather than across ordinary public propagation.
The defect may still reset a session, but the number of unrelated routes and external networks exposed can be limited. Evidence can be captured at every hop.
The limitation is realism. A bounded environment may fail to reproduce topology, policy or software combinations found on the public Internet. That is why staged escalation is preferable: begin with bounded diversity, then widen only when risk and evidence justify it.
Counterfactual 4: downstream routers use narrower UPDATE error handling
Change one fact: a downstream router contains the malformed advertisement without closing the entire BGP session, where a later-style narrow response is applicable.
The experimental route may be discarded, but unrelated valid routes learned over the session remain available. Update amplification and reconvergence pressure should be materially smaller.
This counterfactual reflects the direction later formalized in RFC 7606.[13] It must remain analytical because that RFC postdates the event, and exact handling depends on the error class and implementation.
Counterfactual 5: advance notice reaches affected operators
Change one fact: operators receive sufficient technical notice of the prefix, attribute, window, expected behavior and stop conditions.
Some operators may monitor relevant sessions more closely, prepare staff, constrain exposure or coordinate rapidly after anomalies appear. Diagnosis may accelerate because the route is recognized as an experiment rather than an unexplained event.
Notice does not repair IOS XR. It may also require careful vulnerability handling if a test is expected to expose unsafe behavior. Communication is therefore a mitigation and coordination control, not a complete containment mechanism.
Counterfactual 6: quantitative kill criteria trigger an earlier withdrawal
Change one fact: monitoring identifies abnormal update rates or session resets early enough to cross a predefined stop threshold before 09:08 UTC.
RIPE NCC withdraws sooner. Continued origination ends earlier, potentially reducing repetitions and exposure time.
The limitation is BGP’s distributed state. Updates already propagated would still require withdrawal and reconvergence. Earlier action could reduce duration but would not instantaneously restore all affected paths.
Counterfactual 7: policy limits propagation to selected peers
Change one fact: import and export policies constrain the experimental route to explicitly participating networks.
The failure domain becomes smaller, and participating operators can capture evidence. This aligns with later emphasis on explicit external policy.[14][15]
The limitation is the research question. If the objective requires observing diverse public implementations, strict containment changes what can be learned. That tradeoff should be made explicitly rather than assumed away.
Counterfactual 8: route collectors and service probes provide immediate correlated alarms
Change one fact: control-plane observations, session telemetry and relevant service probes are correlated in real time.
Investigators can distinguish a harmless novel announcement from update amplification, prefix invisibility and service effects sooner. The withdrawal decision becomes evidence-driven.
This does not prevent the first corruption. It improves detection and shortens the interval in which uncertainty persists.
Counterfactual 9: the experiment never occurs
Change one fact: no public announcement is made.
The 27 August trigger disappears, so this event does not expose the defect. The product flaw may nevertheless remain latent and could be activated later by another valid unfamiliar attribute.
This counterfactual clarifies why “do not experiment” is not a sufficient security strategy. Avoiding the test avoids this incident, but it does not correct running software. The better objective is safe discovery: bounded experiments combined with product repair.
10. What verifiable repair would look like
A repair claim should be tied to artifacts and observations rather than reassurance.
Product evidence
For the affected implementation, credible evidence would include:
- the exact corrected software releases;
- a vendor description of the defective handling path at an appropriate level of detail;
- regression tests using valid unknown optional transitive attributes;
- tests showing byte-preserving or otherwise conformant propagation;
- tests using malformed or deliberately corrupted variants;
- evidence that supported narrow error handling preserves unrelated routes where applicable;
- repeated session-cycle tests to detect recurrence;
- operator records identifying installed releases; and
- post-installation observations showing the defect is no longer reproduced.
The public CVE and advisory identify the issue and response.[2][3] They are the beginning of verifiability, not the entire proof of field closure.
Experiment evidence
For a future live routing experiment, credible evidence would include:
- the test prefix and origin;
- the proposed attribute encoding;
- participating peers and intended propagation scope;
- interoperability results across materially different implementations;
- an impact assessment covering control-plane and service consequences;
- an advance-notice record;
- quantitative stop thresholds;
- an immediate withdrawal authority;
- a withdrawal rehearsal;
- live collector monitoring;
- relevant data-plane or service checks;
- timestamps for anomalies and decisions;
- preserved UPDATE data; and
- a post-event account comparing expected and observed behavior.
RIS and Route Views illustrate the value of multiple routing observation points.[7][11] They do not replace device logs or packet captures, but they can independently show whether an announcement propagated, whether withdrawals multiplied and whether effects differed by vantage point.
Operator evidence
An operator asserting that its network is protected should be able to show:
- whether affected IOS XR software is or was present;
- which corrective release is installed;
- how external route policy is defined;
- how malformed UPDATEs are handled by current software;
- how session resets are detected;
- how unrelated route loss is measured;
- which peer protections are enabled;
- how restoration decisions are recorded; and
- whether a controlled regression test has been completed.
This is operational continuity in concrete form. A configuration statement without running-version evidence is incomplete. A software version without policy and observation evidence is also incomplete.
Closure criteria
The event can be considered technically understood when the original bytes, intermediate mutation and downstream response are linked by evidence. The product issue can be considered repaired when fixed software passes relevant regression tests and deployment is demonstrated where required. The governance issue can be considered repaired when a future experiment cannot proceed without documented scope, notice, monitoring, stop authority and retention.
These closure criteria are intentionally separate. A public incident report should state which are satisfied and which remain unknown.
11. Evidence that would alter the conclusion
The current conclusion is evidence-dependent. Several discoveries would require material revision.
Packet captures showing the initiating attribute was malformed
If authoritative captures demonstrated that RIS AS12654 originated an attribute malformed before it reached an affected IOS XR system, the finding that a valid input was first corrupted during propagation would have to change.
Responsibility would shift toward generation and pre-announcement validation, although any additional mutation or amplification would still require separate analysis. The key requirement would be end-to-end byte comparison: what the origin transmitted, what each intermediate system received and what it emitted.
Device logs identifying a different corruption point
If logs or captures showed that another implementation, route server or intermediary created the corruption, the product attribution would need revision. A Cisco advisory can identify a real defect without proving that the same defect explains every observed path.
The event may have contained more than one failure mode. Only path-specific evidence can establish whether all resets shared one corruption point.
Collector data materially revising the impact estimates
If preserved collector data showed that the baseline, affected-prefix count or duration was materially different, the bounded impact assessment should be updated. Corrections might raise or lower the measured scope.
A revision would not automatically change the implementation mechanism. Cause and magnitude are related but independent evidentiary questions.
Approval and risk records showing additional controls
If complete experiment records showed substantial containment, notification or stop controls not visible in the public account, the governance assessment should acknowledge them. It would then need to explain why those controls did not prevent or shorten the observed disturbance.
Conversely, records showing that identified high-impact risks were accepted without mitigation would strengthen the governance criticism. The public account alone does not establish either scenario.
Product records showing prior identification and effective control
If product test records demonstrated that the defect had been identified and effectively controlled before the experiment, the timeline and allocation of responsibility would change. Investigators would need to ask whether the affected deployed systems lacked an available correction, whether operators had received applicable notice and whether the observed behavior came from another mechanism.
The current record does not establish such prior identification.
Evidence of broader or narrower service impact
Complete traffic measurements, operator logs or service telemetry could improve the assessment of user-visible consequences. They might show that control-plane instability caused more application disruption than currently documented, or that redundancy kept most services available despite routing churn.
Such evidence would change the impact section, but it would not justify rewriting the original UPDATE’s validity without packet-level proof.
12. A restrained conclusion
The 2010 RIPE-Duke experiment was not a conventional malicious route hijack, and the available evidence does not support portraying it as one. Nor was it simply a harmless standards test that happened to encounter irrational routers.
It was a live routing experiment in which a valid, unfamiliar optional transitive attribute encountered an affected IOS XR implementation. That implementation corrupted the attribute while propagating it. A downstream router could then respond to the malformed UPDATE by resetting a BGP session, withdrawing unrelated valid routes and contributing to repeated reconvergence. RIPE’s measurements captured a bounded but material disturbance: exceptional update rates, additional prefix invisibility and a peak of nearly 4,500 unstable prefixes.[1][2]
The event exposed two defects in different control domains. One was a product defect in the handling of valid opaque routing information. The other was an experiment-governance weakness: limited pre-public testing and insufficiently bounded public exposure allowed an unknown interaction to become an Internet routing incident.
Cisco’s advisory and maintenance upgrades addressed the product defect. RIPE NCC’s investigation, evidence preservation and stricter future experiment commitments addressed the governance defect. Later standards supplied better error containment and policy guidance, but they are context, not retroactive proof.
The most durable lesson is about control. Standards compliance at origin does not guarantee safe end-to-end behavior. A registry or route collector can identify who announced a prefix and show how visibility changed, but it cannot force every intermediate implementation to preserve an attribute correctly. Vendors must prove safe running code. Experiment initiators must bound uncertain public tests. Operators must know their software, policies and recovery state. Observation systems must preserve enough evidence to distinguish trigger, corruption, amplification and impact.
Accountability is not achieved by naming the first organization in the chronology. It is achieved by matching every consequential control to an owner and requiring evidence that the corresponding repair works.
Sources
- https://labs.ripe.net/author/erik/ripe-ncc-and-duke-university-bgp-experiment/
- https://www.cisco.com/c/en/us/support/docs/csa/cisco-sa-20100827-bgp.html
- https://nvd.nist.gov/vuln/detail/CVE-2010-3035
- https://puck.nether.net/pipermail/cisco-nsp/2010-August/072867.html
- https://www.ripe.net/about-us/executive-board/minutes/2010/minutes-73rd-executive-board-meeting/
- https://seclists.org/nanog/2010/Aug/915
- https://www.ripe.net/analyse/internet-measurements/routing-information-service-ris/
- https://www.ripe.net/analyse/archived-projects/ris-tools-web-interfaces/articles-with-ris-analysis/
- https://stat.ripe.net/AS12654
- https://apps.db.ripe.net/db-web-ui/query?searchtext=AS12654
- https://www.routeviews.org/routeviews/
- https://www.rfc-editor.org/rfc/rfc4271
- https://www.rfc-editor.org/rfc/rfc7606
- https://www.rfc-editor.org/rfc/rfc7454
- https://www.rfc-editor.org/rfc/rfc8212
- https://www.rfc-editor.org/rfc/rfc7908
- https://www.rfc-editor.org/rfc/rfc6811
- https://www.rfc-editor.org/rfc/rfc6480
- https://www.rfc-editor.org/rfc/rfc8205
- https://csrc.nist.gov/pubs/sp/800/189/final
Member Briefing
Deeper Profile Context
Sign in with the right membership level to unlock the full briefing and source notes.
Only for Strategic Circle
Strategic Circle
Open to all readers. Unlock profile briefings after joining and signing in.
Join Strategic CircleOnly for Leadership Alliance
Leadership Alliance
For qualified IP-asset owners and management; sign in to unlock alliance briefings.
Join Leadership Alliance
