Summary
- APNIC Labs’ Table 1 reports 51.73 authoritative queries per
SERVFAILtest and 83.46 when the server stayed silent, compared with 4.40 forNXDOMAIN, 3.93 forNOERROR/NODATA, and 3.43 for a positive control. - The experiment counted what reached its authoritative service. It did not instrument every stub, forwarder, recursive front end, back end, alternate server or transport; the source explicitly leaves the cause of repetition for later study.
- RFC 9520 requires resolvers to cache resolution failures and bound retries, but failover can still move across addresses, servers and transports. Leadership needs a retry budget and an evidence chain, not one aggregate timeout metric.
One intention, several machines that refuse to stop
A browser executes one measurement instruction. It asks for a random name. Somewhere between that act and the authoritative server, the DNS turns the request into A, AAAA and sometimes HTTPS questions. Stub timers may expire. Forwarders may repeat work. Recursive resolvers may choose another authoritative address or transport. A front end may hand the question to a different back end. Each step can be locally reasonable; together they can create a bill no one component believes it authored.
That is the uncomfortable result in Geoff Huston’s 4 September APNIC Blog report. APNIC Labs used scripts embedded in online advertisements to trigger resolutions of randomized names under zones it controlled, then counted queries arriving at its authoritative infrastructure. The randomized label defeated any response cached before that particular test.
The contrast is not subtle. In the report’s Table 1, a positive response averaged 3.43 queries per test. NOERROR/NODATA, which says the requested record type is absent without denying the name itself, averaged 3.93. NXDOMAIN, which says the name does not exist, averaged 4.40. REFUSED averaged 11.47. SERVFAIL rose to 51.73. No response at all reached 83.46.
These are not six equivalent ways to spell failure. A definitive negative gives a resolver reusable information. REFUSED declares a policy decision but says nothing about whether the name exists. SERVFAIL says the responder could not produce an answer and is commonly treated as potentially transient. Silence supplies no DNS statement at all, leaving timers to decide what to try next.
The table is stronger than the prose—and still not a causal trace
The source contains two internal discrepancies worth preserving. Its SERVFAIL paragraph calls 51.29 the average queries per test; Table 1 gives 51.73, while 51.29 is the table’s average repeat count. The no-response paragraph prints the test count as 11,6050,253 and an average of 82.6; Table 1 gives 116,050,253 and 83.46. This report uses the table’s values because that is the source’s explicit summary ledger. It does not pretend the discrepancy never existed.
The ledger is large. There were 138,643,924 SERVFAIL tests and 7,171,673,166 queries, including 6,920,490,780 repeats. Just 3,702,576 tests completed with one query for each type used. The silent condition recorded 9,685,775,212 queries across the table’s 116,050,253 tests; 9,466,474,491 were repeats.
Scale does not create attribution. An authoritative log observes incoming queries, usually from recursive infrastructure. It cannot, on its own, say whether a repeat began at the browser’s stub, a local forwarder, a recursive query distributor, a back-end engine, packet loss, a fresh client event, or a permitted attempt against another address. Huston asks precisely whether recursive implementations or more complex front-end/back-end systems explain the pattern and reserves the answer for follow-up work.
That distinction protects both operators and readers. The measurement proves a distribution of authoritative receipts under controlled response conditions. It does not prove that every SERVFAIL causes 51 queries, that one product is responsible, or that the ad sample can be multiplied into a global traffic invoice.
The baseline already includes legitimate parallel questions
The NXDOMAIN phase ran from 5 to 11 August 2026 over 115,750,503 endpoints recruited through Google Ads. The author says the campaign reached much of the Internet’s geographic and platform diversity, with Russia the major coverage exception.
Not every query after the first was a retry. Forty-eight per cent of endpoints requested both A and AAAA; 51% requested only A; 1% only AAAA. Thirty-nine per cent generated an HTTPS query. If each observed type had been asked once, APNIC calculated about 218 million queries. It received 509,410,787. The 291,919,927-query remainder was classified as repetition. Fifty-seven per cent of endpoints used no more than one query per type; the repeating group averaged 6.03 repeated queries.
The experimental boundary matters. The zone was unsigned, so a validating resolver could not use cached NSEC or NSEC3 proofs to synthesize answers for other random names under RFC 8198. The authoritative service accepted UDP and TCP, not DoT or DoH. Replies were short and untruncated, and arrivals were counted for 24 hours after the ad placement. None of that establishes which transport a client used to reach its recursive resolver.
NOERROR/NODATA performed slightly better than NXDOMAIN in this run: 3.93 versus 4.40 queries per test. That is a result to investigate, not a reason to substitute one response for the other. They assert different facts. RFC 2308 defines how their negative information is cached; RFC 8020 permits a cached NXDOMAIN to stop queries below the nonexistent name. Correct semantics come before lower average traffic.
RFC 9520 draws a retry fence, not a single path
RFC 9520, published in 2023, was written because aggressive resolver behaviour during failures had already amplified incidents. It recounts a Dyn retry storm above ten times normal volume, an 80-fold DNSKEY surge around the root KSK rollover, a 2021 experiment where SERVFAIL lifted one botnet domain’s traffic from about 50 to 60,000 queries per second, and the Facebook outage’s rise from 7,000 to 900,000 queries per second at .COM/.NET infrastructure.
The standard distinguishes useful negative answers from resolution failures. It requires caching failures such as SERVFAIL, unreachable servers and DNSSEC validation failures. The initial cache time should be at least one second and no more than five minutes, with configurable minimums and backoff. A resolver should send no more than two retries to the same server address over the same transport after the original attempt.
But “same address and transport” is a real boundary. Nothing in that limit forbids attempts against other authoritative servers, other addresses or other transports. Nor does it erase repeats triggered farther toward the client. Joining outstanding equivalent queries is therefore as important as counting retries in one loop.
An Extended DNS Error under RFC 8914 can explain that SERVFAIL arose from stale data, DNSSEC trouble, policy or another condition. It is diagnostic context and must not change protocol processing. Better words can improve diagnosis without turning an ambiguous failure into a definitive negative.
A measurement from APNIC is not an APNIC position
The blog identifies Geoff Huston as the author and carries the standard notice that authors’ views do not necessarily reflect APNIC. The institution supplied the publication venue and APNIC Labs measurement system; that does not convert an individual analysis into registry policy.
This separation is especially important because the report is likely to influence resolver and authoritative operators. It is credible evidence for reproducing the test, inspecting retry distributions and checking RFC 9520 behaviour. It is not an instruction to change DNS response semantics for traffic efficiency or a finding that a named operator is non-compliant.
Member Briefing
Deeper Profile Context
Sign in with the right membership level to unlock the full briefing and source notes.
Only for Strategic Circle
Strategic Circle
Open to all readers. Unlock profile briefings after joining and signing in.
Join Strategic CircleOnly for Leadership Alliance
Leadership Alliance
For qualified IP-asset owners and management; sign in to unlock alliance briefings.
Join Leadership Alliance
