Summary
- RFC 2182 recommended three listed servers for most organisation-level zones only as part of a design in which at least one was geographically and topologically removed from the others.
- An advertised server that many resolvers could not reach reduced reliability: distributed probes, retransmissions and timeouts consumed time and network capacity without creating another usable authority.
- Redundancy also depended on maintained state. An NS record, an authoritative reply, a matching SOA serial or a completed transfer each proved a different stage, not identical fresh data everywhere or end-to-end availability.
The zone file showed three; the building had one failure
Three server names appeared in the delegation. On paper, the zone had redundancy. In the machine room, all three shared the same LAN, the same power and the same external link. One local failure could erase the entire set.
RFC 2182 was written for that mismatch. Published in July 1997 as Best Current Practice 16, it did not treat a secondary server as a decorative copy. The memo began from the purpose of multiple authority: zone information should remain widely and reliably available when one server is unavailable or unreachable. A server count was useful only if the servers did not disappear together.
The document therefore required two kinds of separation. Geography protected against a building, room or regional power event. Topology protected against a link, routing segment or service-provider failure. Distance without path diversity could still leave every server behind one cut. Different network names without physical separation could still share one room. Redundancy was the intersection of those facts, not the larger of their labels.
This is why the famous number in the memo was conditional. Two carefully placed servers could sometimes be sufficient, but one prolonged outage would leave only one. RFC 2182 recommended three listed servers for most organisation-level zones, with at least one well removed from the others; four or five could suit a zone demanding higher reliability. It did not say that three NS records certified resilience.
An unreachable server made everyone perform the failure
The memo then moved from placement to advertisement. An NS name should resolve to addresses reachable from the population receiving that information. A machine hidden behind a firewall, an intermittent link or an unusable interface did not become a backup because its address appeared in DNS.
The harm was distributed. A resolver could not know in advance that an address was permanently unreachable. It had to send a query, wait, and often retry because silence looked like ordinary packet loss. An application or user could give up before the resolver exhausted the experiment, turning one unusable address into an apparent failure of the whole zone. Other resolvers would repeat the same work from other networks and at later times.
The 1997 text said that lack of a result was not cached. Errata 4631, held for a future document update, correctly narrows that statement: after RFC 2308, negative answers may be cached, depending on server behaviour and local configuration. The historical detail changed; the operating conclusion did not. Advertising an unusable authority still exports probing, timeout and retry costs to resolvers that cannot repair the topology.
Reachability also had to survive forwarding of the referral. The resolver receiving NS information might pass it to another resolver, so an address usable only from the first vantage point was not enough. RFC 2182 required every address in the returned address RRset to be suitable; operators could not quietly omit a bad address or assign it a special low TTL while presenting the rest as normal. Split internal and external worlds required deliberate DNS views and names, not one misleading global list.
More servers could create less certainty
If two could be too few, why not list twenty? RFC 2182 rejected that shortcut as well. Each extra server enlarged packets, consumed some bandwidth and moved responses nearer to protocol size limits. More importantly, every additional secondary created another configuration, transfer relationship, software instance and administrator that could drift without being noticed.
The memo distinguished listed servers from stealth servers. A site could make several local machines authoritative so local users could still resolve names during an external outage, while advertising only one or two of them globally. Listing every local copy would invite the rest of the Internet to probe the whole site when its common link was down. Local utility did not automatically justify global advertisement.
The result was an optimisation problem, not a quota. Too few independent copies created a brittle service. Too many advertised copies increased coordination and detection costs. The right number followed from the failure domains and the ability to maintain them.
A copy was redundant only while its version remained trustworthy
RFC 2182 ended with a maintenance problem that turns redundancy into time. A secondary used the zone’s SOA serial to decide whether to update. The primary had to increment the serial after every change or group of changes. If an operator accidentally set the serial too high, simply changing it back down could leave a secondary treating the intended correction as older data.
The prescribed repair was staged. The operator advanced through values that were valid increments under RFC 1982 serial arithmetic, waited for every relevant secondary to adopt each stage, and only then crossed the wraparound boundary to the desired lower absolute number. The memo’s strongest instruction was not the example values. It was to verify every stage and assume nothing.
That sequence exposes several distinct receipts. A listed NS record proves that a name was advertised. A response proves that one server process answered. An SOA query proves the serial that one vantage point observed at one moment. A zone transfer proves that a transfer procedure completed. None by itself proves that the loaded zone is byte-for-byte the intended version, that every serving process switched atomically, that the server sits in an independent failure domain, or that an application can reach its final service.
The security section was equally bounded. The BCP did not claim to add or remove DNS security problems, but warned that compromise of a secondary could in some circumstances compromise hosts in the domain. Operational independence increases availability only when the added operator and machine remain trustworthy enough to serve authoritative data.
RFC 2182’s lasting contribution was to move DNS redundancy from inventory to evidence. Count the servers, but also map what can fail together. Publish addresses, but test them from the networks expected to use them. Replicate the zone, but observe version progression and serving state after change. The zone file describes authority; running systems have to earn availability.
Sources
Member Briefing
Deeper Profile Context
Sign in with the right membership level to unlock the full briefing and source notes.
Only for Strategic Circle
Strategic Circle
Open to all readers. Unlock profile briefings after joining and signing in.
Join Strategic CircleOnly for Leadership Alliance
Leadership Alliance
For qualified IP-asset owners and management; sign in to unlock alliance briefings.
Join Leadership Alliance

