Summary

  • APNIC’s one-nameserver run covered 254,894,985 tests and averaged 3.43 authoritative queries per test; its later two-nameserver run covered 150,221,951 tests and averaged 2.57.
  • The share of tests producing exactly one query rose from 58% to 71%, but the runs were separate and unequal. The published aggregates do not establish that adding a nameserver caused the change.
  • The useful next artefact is a reproduction receipt: common windows, randomized treatments, comparable resolver and network cohorts, frozen authoritative topology, a published normalization formula and timing distributions that test alternative mechanisms.

A striking difference between two baselines

APNIC Labs asked a practical question: does offering two authoritative nameservers change how many queries recursive resolvers send? The published answer was surprising. In the earlier one-nameserver dataset, 254,894,985 tests generated 875,316,423 authoritative queries—an average of 3.43 per test. Some 147,233,117 tests, reported as 58%, generated only one query.

The later two-nameserver dataset contained 150,221,951 tests and 385,725,364 queries, an average of 2.57. Here 106,376,529 tests, or 71%, generated only one query. Among tests with repeats, the reported average number of repeated queries also fell, from 3.80 in the first run to 2.56 in the second.

Those are large datasets and material aggregate differences. They support a clean descriptive statement: the later two-nameserver run contained fewer observed queries per test and more single-query tests. They do not, on their own, identify the extra nameserver as the cause. The source itself calls the result unexpected and says it cannot offer a definitive explanation.

Sample size is not the same as comparability

The unequal totals—about 255 million tests versus 150 million—do not automatically invalidate a comparison. Large observational datasets are routinely normalized. APNIC did normalize the two datasets before comparing their timing profiles. The harder question is whether the two samples represent sufficiently similar measurement conditions.

The preceding experiment describes unique random names delivered through Google advertising endpoints, short untruncated answers from an unsigned zone, UDP and TCP service, and a 24-hour observation window. It also reports geographic effects, with Russia as a notable exception in reach. The one-server setup placed one dual-stack authoritative server in each of six broad zones: North America, South America, Europe plus Africa, India, Asia and China.

For a causal comparison, a reader needs to know what was held constant when the second nameserver was added. Were the treatments run in the same dates and hours? Did the mix of countries, access networks, resolver addresses and address families remain stable? Were routing, server placement, anycast or unicast design, TTLs, response codes, packet sizes and retry observation windows identical? Aggregate division can align denominators; it cannot align an unreported change in cohort or network path.

This is why the right criticism is not “the samples are different sizes.” It is “the treatment receipt is incomplete.” A smaller second run can still be persuasive if assignment and controls are visible. An enormous run can still leave causation unresolved if treatment arrives with a different clock or population.

The timing shapes contain the better clue

The experiment’s distributions are more informative than the headline averages. In the one-nameserver run, more than 85% of duplicates arrived within the first second. With two nameservers, the corresponding share was 75%. In both runs, about 90% arrived within five seconds. The two-server curve also showed peaks near 0.75, 1.5 and 3 seconds.

At the fastest end, 17% of repeats in the one-server data arrived within 10 milliseconds, compared with 12% in the two-server data. After normalization, the largest difference was concentrated between 10 and 70 milliseconds, especially 10 to 40 milliseconds. That is shorter than a typical recursive resolver’s UDP timeout. It therefore invites mechanisms other than a simple “wait, fail, retry” story.

One possibility raised by the author is a resolver farm: a DNS-aware front end can distribute work across backend recursive resolvers, and more than one backend might pursue the same name. PowerDNS describes dnsdist as a DNS-aware load balancer placed in front of recursive servers. That establishes that such an architecture exists, not that it produced APNIC’s curves. Without internal traces, the observed source address may hide frontend fan-out, backend races, address-family choices, cache state or implementation-specific nameserver selection.

Resolver software also does not choose authoritative servers through one universal algorithm. APNIC’s notes from DNS-OARC describe different initial scores, latency learning and IPv6 preferences in BIND, PowerDNS Recursor, Knot Resolver and Unbound. A second dual-stack name adds choices of names and addresses; its effect depends on implementation and topology. “Two nameservers” is therefore not a complete experimental variable.

The reproduction receipt

A decision-grade follow-up can be compact, but it must bind the two treatments to the same operating conditions.

First, assign tests randomly between one- and two-nameserver configurations during the same measurement window, or alternate treatments in short, predeclared blocks. Publish test counts for each block. This prevents a day, campaign or traffic shift from silently becoming the treatment effect.

Second, publish cohort composition by broad region, access network, stable resolver address and address family, using privacy-preserving aggregates. The purpose is not to identify users. It is to show whether the 58% and 71% populations are comparable and whether the aggregate difference survives stratification.

Third, freeze the authoritative surface except for the added name: server software, zone content, TTL, DNSSEC state, response code, response size, truncation policy, UDP/TCP handling, observation period, addresses, locations and routing. State whether each endpoint is unicast or anycast. Two labels that reach the same facility do not exercise the same resilience or selection problem as two topologically diverse services.

Fourth, publish the normalization formula, exclusions and denominators, then release raw and normalized distributions with uncertainty intervals. Report effects separately for sub-10ms bursts, the 10–70ms region, later timer peaks and the tail beyond five seconds. Averages conceal precisely the shapes that distinguish fan-out, racing and timeout hypotheses.

Finally, predeclare alternative explanations. Frontend fan-out, backend concurrency, IPv4/IPv6 selection, recursive implementation, cache state and campaign geography should compete with the server-count explanation. A mechanism earns confidence when it predicts a timing or cohort pattern that the reproduction then observes.

Two nameservers remain sensible for a different reason

Nothing in this audit reverses the operational case for multiple authoritative nameservers. RFC 2182 recommends multiple, physically and topologically separated servers for reliability and load spreading. APNIC’s earlier resilience study likewise suggested at least two distinct dual-stack, geographically diverse services, with diminishing marginal gain beyond two in its measurements.

That guidance stands on continuity: one server, path or facility can fail while another remains reachable. It should not be made to depend on an unproven promise that adding a second nameserver will reduce aggregate recursive queries from 3.43 to 2.57. Conflating the two claims weakens both. Operators can follow the resilience rule now and still demand a controlled reproduction before using the APNIC difference for capacity models, resolver design or performance forecasting.

The disciplined conclusion is narrow and useful. APNIC observed a substantial shift between two published baselines. The source did not pretend to know why. A shared-clock reproduction receipt would turn that honest uncertainty into a testable operating result.

Sources