Summary
- RIPE-400 reported that roughly 11–13% of name servers in a March 2006 delegation survey were lame under its test, but that configuration share was not a query-weighted measure of user harm.
- By 2009, RIPE NCC had evidence that the test and the notification process needed refinement. A traffic study estimated that about 1% of observed queries were affected, with important methodological limits.
- The 2009 annual report says periodic lameness checks continued while email alerts stopped. The durable lesson is to measure detection, delivery, attention, repair and user outcome as separate stages.
The switch that did not turn off the measurement
An email programme can end while the underlying observation continues. RIPE NCC's 2009 annual report records exactly that split: the organisation continued periodic lameness checks, but decided to stop sending email alerts. That sentence is the control point for reading the preceding three years. It prevents an attractive but false account in which a registry first measured broken DNS delegations and then abandoned the measurement when operators did not respond.
The historical record describes something more demanding. RIPE NCC's operational experience put a measurement, a contact route and an intervention under scrutiny at the same time. Each could succeed or fail independently; the record does not provide a controlled experiment attributing repairs or user outcomes to the emails.
RIPE-400, published in January 2007, began with a March 2006 survey. Around 11–13% of name servers in delegations from IANA to RIPE NCC were classified as lame. The test was not merely “did a host answer?” The target name had to resolve to an IPv4 or IPv6 address, and a recursion-disabled UDP query for the zone's SOA record sent to that same address had to return a single authoritative SOA answer. Failed cases were retried five times over ten days. The proposed service would run monthly, derive contacts from SOA RNAME and maintainer data, send one message per lame server, publish statistics and periodically assess effectiveness.
That was a serious attempt to turn a configuration observation into operational maintenance. It was not yet proof that every flagged record imposed equal cost, that every message reached the controlling operator, or that a reply or repair improved a user's lookup.
Eleven per cent of records is not eleven per cent of experience
The most important correction came from asking a different denominator question.
A delegation survey counts tested records. Users generate queries. A server can be configured badly yet receive almost no traffic because other authoritative servers answer first or caches absorb the demand. Another defective server can sit on a heavily queried path and create repeated delays. Counting both as one defective record is useful for inventory hygiene. It is not a measurement of equal consequence.
The April 2009 DNS Working Group debate made this distinction explicit. The existing method did not separate a timeout from every other reply that failed the authoritative-answer rule. Operationally, those are different states. A timeout may make a resolver wait and retry another server. A prompt, clear non-authoritative response may let it move on sooner. Both can indicate a delegation inconsistency, but they do not impose the same latency, retry load or diagnostic path.
The critique also challenged the easy interpretation of the 11–13% figure. The user effect could be greater or smaller depending on how queries were distributed. No measured user-impact number followed automatically from the configuration percentage.
At RIPE 59, the analysis moved closer to consequence. The “Falling Trees” presentation compared lame-zone findings with an hour of traffic observed at a reverse-DNS master. From more than 16 million scanned packets, it estimated roughly 0.3% bad NS records, about 0.8% of A/AAAA conditions causing DNS lookup failure, and about 1% of observed queries affected. Those figures were not a universal rate. The presentation documented limits: it did not model caching, and parts of the classification operated at IP level rather than IP-plus-domain level. Its value was methodological.
It tested prevalence against use instead of assuming the two were interchangeable.
The alert chain had several breaks
The notification experiment also exposed a chain of evidence that a simple “email sent” counter could not supply.
First came detection: was the observed condition a stable fault, a transient timeout, a deliberate configuration, or a probe artefact? A February 2009 update said replies to a small batch sent in October 2008 had revealed problems in probes and result interpretation. RIPE NCC refined the system before starting more small batches on 26 February.
Second came delivery: did SOA RNAME or maintainer data lead to a monitored address, and did that address belong to someone who controlled the delegation? The presence of a contact field does not prove receipt by an accountable operator.
Third came attention: did a technically valid notice deserve priority among maintenance work? The RIPE 59 material observed that lame servers with no apparent use tended not to be fixed, while used ones were fixed. That is consistent with operators allocating attention by consequence. It is not causal proof that a particular email produced a repair.
Fourth came verification. At RIPE 58, participants asked how an operator could confirm that a change had cured the condition. Related reverse-DNS tools existed, but their tests differed subtly. A control loop that warns without offering an exact re-test leaves the recipient uncertain about closure.
Finally came outcome. A changed delegation record is an intermediate state. The relevant public result is whether timeouts, retries or failed lookups decline on paths that users actually exercise.
The working-group discussion therefore became less about whether lame delegations were desirable—they were not—and more about the proportional response. Repeated unsolicited notices impose costs on recipients and sender. Minutes from RIPE 59 record concern that continuing to contact people who were actively ignoring the messages could resemble harassment. An October proposal recommended ending mass mail, providing an annual report to LIRs and targeting the most consequential cases.
A stronger “clean data” position, including escalation toward pulling a delegation, appeared in debate; the reviewed record does not establish it as the adopted policy.
A better way to prove that repair is worth the interruption
The historical programme suggests an evaluation design that is still useful for infrastructure operators.
Start by classifying the defect rather than collapsing every failed check into “lame”. Keep at least timeouts, explicit non-authoritative replies, address-resolution failures and inconsistent authority as separate cohorts. Record repeat observations so a transient response is not mistaken for a persistent state.
Then estimate exposure. Join the configuration finding to privacy-preserving, bounded traffic evidence where that is lawful and technically appropriate. The objective is not to rank operators publicly. It is to distinguish an unused inconsistency from a failure mode that produces material retry time or lookup loss.
Next, instrument the intervention. Record whether a notice was generated, delivered, acknowledged, assigned, changed and re-tested. These are separate events. A registry should not claim success from the first event when the claimed benefit lies in the last.
Finally, construct a comparison. Stagger contact by severity, retain a suitable holdout where operational ethics allow, or compare similarly exposed cohorts receiving different messages. Measure time to repair, recurrence and query-weighted outcome. A before-and-after fall without a comparison can reflect unrelated operator work, seasonal traffic or a changed probe.
This changes the purpose of notification. The email is no longer the product. It is one intervention inside a measurable control loop. Low-impact cases can be surfaced in an annual report or self-service dashboard. High-impact, well-diagnosed cases can receive targeted contact with an exact reproduction and verification path. If neither route changes outcomes, the programme has evidence to redesign itself rather than merely send more messages.
The current RIPE NCC documentation reviewed for this article describes name-server checks when reverse-DNS delegations are created or changed. It is not evidence about the present status of the older periodic programme or its email policy. The claim here remains historical and narrower: in 2009, the checks and the emails were separable, and RIPE NCC said it kept the former while stopping the latter.
Sources
- RIPE-400: RIPE NCC DNS Lameness Checking
- RIPE-400 official PDF
- DNS Working Group archive, February 2009
- DNS Working Group technical critique, April 2009
- RIPE 58 DNS Working Group minutes
- Falling Trees: lameness and traffic at RIPE 59
- RIPE 59 DNS Working Group minutes
- Solving DNS Lameness discussion, October 2009
- RIPE NCC Annual Report 2009
- Current RIPE Database reverse-DNS configuration documentation
Member Briefing
Deeper Profile Context
Sign in with the right membership level to unlock the full briefing and source notes.
Only for Strategic Circle
Strategic Circle
Open to all readers. Unlock profile briefings after joining and signing in.
Join Strategic CircleOnly for Leadership Alliance
Leadership Alliance
For qualified IP-asset owners and management; sign in to unlock alliance briefings.
Join Leadership Alliance
