Summary
- Route flap damping answered a real early-Internet constraint: repeated BGP changes could consume scarce router processing capacity and contribute to further instability.
- Its penalty counter could not distinguish chronic physical failure from the alternate-path exploration that follows an ordinary withdrawal, so a repaired and valid prefix could remain suppressed in networks its holder did not control.
- The durable authority did not sit in one standards body. RFC authors described the mechanism, vendors encoded defaults, and network operators decided whether those defaults would execute.
- After research exposed the convergence penalty, the operational path moved from coordinated settings to broad disablement and then to more conservative thresholds. A deeper selective algorithm did not become the common repair documented by the later standards.
Three changes in ten minutes
The maintenance record in RIPE-229 begins with a scheduled software upgrade. The reload after the upgrade counted as one route flap. The new software crashed, adding another. Reloading the old software added a third. Three changes occurred within ten minutes.
The backbone had not entered an endless cycle. Its operator had tried a change, found it bad and restored the previous state. Yet some upstream and peering routers applied aggressive, prefix-length-sensitive damping. The document records that the operator's /24 routes, including customer routes, were held down for more than three hours at some boundaries. Repair had produced the signals of misconduct, and the remote counters did not know the difference.
That episode helped shape the coordinated parameter recommendation. At RIPE 27, the working group chose not to begin suppression before a fourth consecutive flap and to cap suppression at one hour after the last flap. The limit was gentler than the behavior just observed. It was still a rule under which a working /24 could disappear for an hour after crossing the threshold.
The problem RFD was built to solve
Route flap damping was not invented in search of administrative power. Its early documents describe a control-plane problem. In the first half of the 1990s, the number of announced prefixes was rising, inter-provider paths were becoming denser, and every withdrawal could be processed by routers carrying a full table. Early routing engines had far less processing headroom than modern equipment. A burst of updates could load a router badly enough that it lost sessions; that failure could send still more changes to its neighbours.
Operational documents date the mechanism to 1993 and its appearance in Cisco, ISI/RSd and GateD software to 1995. By the time RFC 2439 was published in November 1998, its authors described the technique as implemented in commercial products and widely deployed. The RFC's goals were precise: reduce processing load, prevent sustained oscillation and avoid delaying generally well-behaved routes.
The mechanism kept a score for a route learned from an external BGP neighbour. A withdrawal raised the score. Other changes could raise it too, depending on the implementation. The score decayed exponentially with time. Above a suppress threshold, the route would not be used; below a reuse threshold, it could return. A maximum hold-down bounded the memory.
This was an efficient answer to limited hardware. It was also a prediction system. RFC 2439 acknowledged that future route stability could not be known accurately, so recent changes served as a proxy. The crucial question became what counted as a change and how much punishment its history should carry.
From configurable mechanism to coordinated rule
Local choice created a second problem. Vendors shipped different defaults. Operators could select different half-lives, thresholds and maximum suppression times. A route might be available through one part of the Internet while still suppressed through another. The origin network could fix the fault and clear its own router, yet have no direct way to clear a counter in an unknown upstream autonomous system.
RIPE working-group documents tried to make those distributed decisions predictable. A discussion at RIPE 26 led to a BOF; RIPE 27 produced a task force; RIPE-178 appeared in 1998. Tony Barber of UUNET designed a parameter set that had run in his environment for several months. RIPE-229 refined the recommendation in 2001 and urged its use by ISPs and as vendor defaults.
This was influence, not command. RIPE could publish a common setting but could not log into every router. RFC 2439 could standardise a mechanism but could not enable it. The executable authority stayed with vendors and operators: vendors decided which counters and constants existed, while an operator decided what ran on equipment it controlled.
The recommendation also made a distributive choice. Its “graded” approach treated /24 and longer prefixes more harshly than shorter aggregates. After the threshold, a /24 could receive a one-hour outage; shorter prefixes received lower ranges. The stated reasoning included aggregation and the larger number of users behind shorter prefixes. But multihomed small networks often needed longer, unaggregatable announcements. The burden therefore fell unevenly on holders whose routing resilience depended on specificity.
The exception revealed the rule. RIPE-229 proposed “Golden Networks” for root and top-level-domain name servers that happened to sit in longer prefixes. Critical infrastructure needed protection from the same policy intended to protect infrastructure. The list did not make prefix length a better measure of future behavior; it identified prefixes whose false suppression was considered too costly.
When one withdrawal looked like many failures
The decisive challenge came from the behavior of BGP itself. After a route is withdrawn, neighbouring routers do not necessarily move directly from one valid path to “unreachable.” They can explore a sequence of alternate paths. Each best-path change carries a different AS path. A damping implementation observing those attribute changes can add penalties even though the origin has failed only once.
In their 2002 SIGCOMM paper, Zhuoqing Morley Mao, Ramesh Govindan, George Varghese and Randy Katz analysed this interaction with a model, simulations, traces and a commercial-router testbed. In studied topologies, a single withdrawal followed by a reannouncement could produce enough secondary changes to suppress the returning route for up to an hour. RIPE-378 later described a measurement in which one prefix withdrawal produced 41 BGP events a few autonomous-system hops away.
The finding was bounded. The paper did not measure how much of the world ran RFD, and it left the exact Internet topologies that would trigger the effect for later work. But it falsified an important assumption behind aggressive settings. A high score did not necessarily mean repeated physical instability. It could be the routing system counting its own search for an alternative.
The authors proposed selective flap damping: do not count monotonic path changes characteristic of withdrawal exploration as fresh flaps. In the topologies they tested, the change eliminated withdrawal-triggered suppression while still damping a route that genuinely alternated every 40 seconds. It was a plausible redesign, not proof of production-scale adoption.
The cure was first removed, then retuned
RIPE-378 recorded the operational response in 2006. It said that no demand from the ISP industry and no implementer activity had emerged for the 2002 modifications. Router processing power had also increased, reducing the urgency of the problem RFD first addressed. The document declared the earlier parameter recommendations obsolete and advised against using then-current RFD implementations in ISP networks. Its authors concluded that the reachability side effects were likely worse than simply running without damping.
A 2012 Internet-Draft survey provides a limited view of what followed. It gathered 63 self-selected responses after calls to operator mailing lists. Thirteen respondents said they used RFD; 49 said they did not. Fifteen answers cited RIPE-378 as a reason for non-use. The sample cannot establish a worldwide deployment rate, but it supports the narrower claim that the recommendation had become operationally legible and that customer impact was a reported concern.
The later return did not erase the 2002 mechanism. It narrowed whom the counter would catch. Measurement behind RIPE-580 and RFC 7196 found that a small share of prefixes generated a large share of updates. In the week studied for RFC 7196, raising the suppress threshold from 2,000 to 6,000 reduced the update rate by 19 per cent compared with no damping while damping 90 per cent fewer prefixes than the 2,000 setting. At 12,000, only 0.22 per cent of prefixes were damped in that experiment, with about an 11 per cent reduction in average hourly updates.
RFC 7196 therefore recommended at least 6,000 for less destructive but still aggressive use and at least 12,000 for a conservative setting. It also suggested an observe-only mode in which an operator could calculate the effect without suppressing routes. Yet it advised implementations not to change existing defaults silently, because doing so could break operational configurations. A verified erratum makes clear that the Cisco and Juniper defaults listed in its table were informational, not endorsed values.
This is how lock-in appears without a central ruler. The old defaults were judged too aggressive, but changing them automatically carried its own compatibility risk. The selective algorithm required code and deployment work; a higher threshold fit an installed configuration surface. Retuning was the path of lower coordination cost.
Who controlled reachability, and who paid
The prefix holder controlled its own announcement and repair. A registry record could show which organisation held the address resource. Neither fact forced a remote router to reuse the route. Transit and peering operators controlled damping on their equipment. Vendors controlled which knobs existed and how those knobs behaved if left alone. Standards and working-group authors supplied a language, a mechanism and recommended values, but no document executed a configuration.
The immediate beneficiary was the network protecting its control plane from repeated updates. In severe early cases, everyone could benefit if that protection prevented a router failure from cascading. The cost in a false or amplified case landed elsewhere: the holder of a repaired prefix, its customers and users trying to reach its services. Diagnosis could span several autonomous systems, while the affected party might not know which one still held a penalty.
That arrangement was locally authorised in a narrow sense: an operator configured a router it owned. It was not a global mandate granted by every affected prefix holder. Nor does the evidence prove a contract breach or unlawful act. The legitimacy problem is more specific. A consequential decision was hard for the affected party to observe, attribute or contest.
The counterfactual cannot delete the early constraint
An Internet without RFD in 1995 would not automatically have been faster and safer. The early processing problem was real in the records, and no-damping could have meant more churn and more overloaded routers. A credible alternative needs another way to protect control planes.
Selective damping is one such counterfactual, but the 2002 paper proves only what occurred in its studied environments. Conservative thresholds from the beginning are another. They would probably have spared more ordinary prefixes while allowing more updates, the trade-off measured later. Better visibility is a third: expose the penalty, the observing AS and a safe clearing process to the affected network. That could shorten diagnosis without creating a central routing authority, although authentication and abuse would remain difficult.
The historical result is less dramatic and more durable. A standard described. Software encoded. Operators executed. When those powers were collapsed into “the Internet's stability policy,” the person paying for delay disappeared from view.
For number-resource holders, RFD supplies a hard boundary. A registry can record whose prefix it is, and an origin can announce it correctly, while remote networks still decide whether the route is believable now. Resource title is evidence. Reachability is a continuing relationship among autonomous systems.
Sources
- RFC 2439: BGP Route Flap Damping
- RIPE-178: early route-flap damping parameter recommendation
- RIPE-229: Coordinated Route-flap Damping Parameters
- Mao et al., Route Flap Damping Exacerbates Internet Routing Convergence
- RIPE-378: Recommendations on Route-flap Damping
- 2012 Route Flap Damping Deployment Status Survey
- RIPE-580 comparison and revised recommendation
- RFC 7196: Making Route Flap Damping Usable
- Verified erratum 4011 for RFC 7196
Member Briefing
Deeper Profile Context
Sign in with the right membership level to unlock the full briefing and source notes.
Only for Strategic Circle
Strategic Circle
Open to all readers. Unlock profile briefings after joining and signing in.
Join Strategic CircleOnly for Leadership Alliance
Leadership Alliance
For qualified IP-asset owners and management; sign in to unlock alliance briefings.
Join Leadership Alliance
