Summary
- George Michaelson wrote that his home DNS backup existed only in theory; when he rebooted the supposed primary, he discovered it was the only instance.
- DHCP and Router Advertisement standards can carry several resolver addresses, but an advertised list does not prove that clients retained an alternative, used it during failure or received the same local DNS policy.
- A proportionate failover receipt should connect six separate states: running endpoint, independent failure domain, issued configuration, client observation, bounded failover and tested rollback.
On 9 September, George Michaelson published a short account on the APNIC Blog about running Pi-hole and AdGuard as in-house DNS services. It is not an APNIC production incident notice; the page expressly presents the views as the author’s own. Its value lies in a sentence that strips a familiar word of false comfort. Michaelson thought he had a backup plan. He had not implemented it. When he rebooted what he regarded as the primary DNS service, devices around the house began to fail and the “primary” turned out to be the only instance.
That is a small event with a clean evidentiary boundary. The post supplies no packet trace, outage duration, device inventory, topology or security finding. It does show the difference between intention and operating state. A diagram can contain a secondary resolver before any process is running. A router can contain two DNS addresses before a client has received them. A client can retain two addresses without ever selecting the second one. The second endpoint can answer ordinary names while silently lacking the first endpoint’s blocking or redirection rules. Each transition needs its own observation.
The outage began before the reboot
The visible interruption began when the DNS service was restarted. The resilience failure began earlier, when the backup remained a plan. That distinction matters because change control is usually organised around a belief about the current system. If an administrator believes a tested alternative exists, rebooting one endpoint looks like a bounded operation. If no alternative is running, the same action becomes a network-wide dependency test performed on live users.
Michaelson’s description makes the dependency surface explicit. DNS reboots stopped other household functions. Router changes could interrupt wireless service. Changing the router’s DNS configuration affected DHCP and produced what he called a freeze-and-thaw cycle. His three questions therefore reach beyond server count: which devices depend on the service, whether the network will actually adopt an alternative, and whether the old state can be restored.
Those are better questions than “Do I have two boxes?” Two resolver processes on one host share a host failure. Two hosts on one power supply share a power failure. Two addresses distributed by one router share its configuration and availability. Two installations copied from the same mistaken configuration share an administrative failure. Independence is not a decorative adjective; it is a claim about which failure can remove both paths at once.
A server list is an instruction, not an outcome
The standards are useful precisely because they show where configuration evidence stops. RFC 2132’s DHCPv4 option 6 can present a list of DNS servers and says they should be listed in preference order. RFC 3646 does the corresponding job for DHCPv6, allowing one or more IPv6 recursive-server addresses in an ordered list. RFC 8106 permits IPv6 Router Advertisements to carry one or more Recursive DNS Server, or RDNSS, addresses and attaches a lifetime to them; a lifetime of zero withdraws the addresses from use.
These are provisioning mechanisms. They can prove what a router or DHCP service intended to issue at a particular time. They cannot, by themselves, prove what a television, phone, laptop, sensor or application stored, preferred or queried. RFC 6419 documented little commonality in how desktop systems handled multiple DNS lists and interfaces: some kept per-interface lists, others one system-wide list, and selection or backoff differed. The document is not a current survey of Michaelson’s household, but it is enough to defeat the assumption that one router-side screenshot describes every client.
RFC 9520 supplies another useful boundary. A stub resolver may be configured with multiple recursive resolver addresses, and a resolver is not prevented from retrying at another server or over another transport. Yet the RFC does not promise that a particular client will do so on a particular timetable. It defines resolution failure only after none of the available servers supplies useful data for the query. Between “two addresses were advertised” and “the user still received a useful answer” lies the behaviour that a failover test must observe.
Reachability can survive while policy fails
Michaelson was not merely forwarding arbitrary DNS queries. He says he used Pi-hole and AdGuard and redirected some known names to unrooted IP addresses. That makes policy parity part of the resilience claim. Suppose the preferred resolver blocks an advertising domain or supplies a local answer, while the alternative returns the public answer. The household may remain online during failover and still lose the behaviour that justified self-hosting.
This does not make a difference in answers a security breach. The source contains no such finding. It means that a useful failover receipt should ask two questions, not one: did resolution continue, and did the alternative provide the intended local policy? Availability without policy consistency is a partial success that should be recorded as such.
The right evidence can remain small
The strongest objection is proportional. Most homes can tolerate occasional outages; Michaelson says so himself. Running two resolvers, keeping rules aligned, separating power or host dependencies and testing several device classes may cost more effort than a brief interruption. Turning a household lesson into enterprise ceremony would miss the point.
The answer is not a service-level agreement. It is a short receipt made before the next maintenance window. Record resolver A and B, where they run, and which power, switch, router and upstream links they share. Capture the DNS addresses and lifetimes being issued through DHCPv4, DHCPv6 or RDNSS. Check what a few materially different clients retained. Withdraw the preferred resolver briefly. Ask a normal name and a name whose local treatment matters. Where observable, record which resolver answered and how long useful service took to return. Then restore the preferred path and confirm the rollback.
That record does not prove every device, query type, transport or outage mode. It does something more honest: it closes a bounded case and names what remains untested. A plan becomes a backup when the alternative is executable. It becomes resilience when clients can use it under failure. It becomes maintainable when the previous state can be restored.
Sources
- George Michaelson, “Building resilient self hosted services is not always easy”, APNIC Blog, 9 September 2026
- RFC 2132: DHCP Options and BOOTP Vendor Extensions
- RFC 3646: DNS Configuration options for DHCPv6
- RFC 8106: IPv6 Router Advertisement Options for DNS Configuration
- RFC 6419: Current Practices for Multiple-Interface Hosts
- RFC 9520: Negative Caching of DNS Resolution Failures
Member Briefing
Deeper Profile Context
Sign in with the right membership level to unlock the full briefing and source notes.
Only for Strategic Circle
Strategic Circle
Open to all readers. Unlock profile briefings after joining and signing in.
Join Strategic CircleOnly for Leadership Alliance
Leadership Alliance
For qualified IP-asset owners and management; sign in to unlock alliance briefings.
Join Leadership Alliance
