Summary

  • RFC 8767 does not abolish DNS TTL. Expiry still requires the recursive resolver to consult the source again; stale use begins only when authoritative refresh cannot be completed and an eligible older cache entry exists.
  • Four different timers govern the exception: the client's useful response window, the resolver's total resolution work, the recheck interval after failure and the maximum age of retained stale data. A single “serve stale enabled” flag proves none of them.
  • The evidence record must keep the old RRset, refresh failure, DNSSEC time, returned short TTL, optional Extended DNS Error, continued refresh and application outcome separate. Availability is a local fallback decision, not authority to declare old data current.

The address that survived its owner

Imagine a payment gateway moving from one hosting provider to another. Its operator publishes a new A record and keeps the old service alive during the DNS transition. Most recursive resolvers refresh and follow the move. One enterprise resolver, however, cannot reach any of the zone's authoritative servers during a network incident. It still holds the old address, whose TTL expired forty minutes earlier.

There are two unattractive outcomes. Returning failure interrupts every user behind that resolver even though the old endpoint may still work. Returning the old address preserves service if the overlap remains, but could send credentials or transactions to an address that has already changed hands. The packet does not reveal which world exists. The resolver only knows what it previously received, when ordinary reuse ended, the observed refresh outcome and which local exception its operator authorized.

That is the real serve-stale problem. It is not “freshness versus uptime” in the abstract. It is custody of an old statement after the authority that issued it can no longer be consulted. The resolver can make failure less destructive, but it cannot manufacture current knowledge.

TTL expiry is a demand to ask again

RFC 1035 described TTL as the interval for which an RR may be cached before the source should again be consulted and before the record should be discarded. RFC 2181 emphasized a maximum time to live. RFC 8767 updates that model. Its revised language makes the distinction precise: the record may be cached for the TTL, and when that duration ends the source must again be consulted; if authoritative refresh is unavailable, the record may be used as though unexpired under the document's conditions.

The order matters. A resolver does not acquire a general right to ignore the publisher's TTL. Expiry ends ordinary cache authority and starts a refresh obligation. Serve-stale is the exceptional branch after that obligation has produced no usable authoritative replacement within the applicable time budget.

Zero is a hard boundary. An RR with TTL 0 is for the transaction in progress and is not to be cached. It cannot later be promoted into stale fallback. The original TTL clamp is also not the stale window. RFC 8767 recommends capping received TTLs around seven days, while its separate maximum-stale timer governs how long an already expired object remains eligible. Conflating the two hides both the publisher's freshness instruction and the resolver's emergency extension.

One feature, four clocks

The RFC deliberately gives an example method rather than one universal algorithm. Its four timers describe different authorities.

The client response timer asks how long the resolver can attempt a normal refresh before the answer becomes useless to the waiting client. The example suggests about 1.8 seconds, just below a common two-second timeout. A shorter value improves apparent latency but may serve old data when an authority is merely slow. A longer value protects freshness but may deliver only a late failure.

The query resolution timer bounds the total upstream work. It is commonly much longer, around ten to thirty seconds in the RFC's discussion. Crucially, the resolver should keep working after it sends a stale answer. The client deadline and the repair deadline are not the same event.

The failure recheck timer prevents every client query from hammering an already failing authority. A recent failed refresh can license immediate reuse of stale data until recheck is due. Too short a period magnifies an outage or attack; too long a period hides recovery. RFC 8767 recommends attempts no more frequently than every thirty seconds in its example and points to a five-minute ceiling. RFC 9520 now requires bounded caching of resolution failures and recommends backoff.

The maximum stale timer determines how long after ordinary expiry the resolver retains an answer as eligible emergency memory. The example suggests one to three days. A shorter window loses resilience during a long authority outage; a longer one consumes memory and enlarges the interval in which abandoned or deliberately changed data can still act.

These clocks cannot be compressed into one dashboard badge. A resolver might retain stale entries but never return them. It might return them immediately without first waiting. It might keep them for one day but suppress upstream rechecks for five minutes. It might behave differently after a restart or under cache pressure. Running configuration has to be observed at each clock boundary.

What counts as a refresh

Not every packet from an authoritative address replaces the old state. RFC 8767 says an authoritative answer with the AA bit and either NOERROR or NXDOMAIN must count as refresh. Other response codes should normally count as failure, leaving the previous cache state intact.

This rule prevents a transient SERVFAIL from erasing the last useful answer. It also creates an important asymmetry. A fresh authoritative NXDOMAIN can legitimately terminate a previously positive answer, even when an operator suspects that deletion was accidental. The resolver has no reliable way to distinguish a deliberate removal from an erroneous one. Stale service is not a veto over a current authoritative denial.

REFUSED is harder because it carries several meanings. It may reflect an explicit takedown intention, a temporary ACL mistake or a server that is no longer authoritative. RFC 8767 permits implementers to decide whether all authorities returning REFUSED should preempt stale use. That choice belongs in the local semantic matrix and incident runbook; it must not be inferred from the option name.

RFC 9520 adds another state that operators must not confuse with an answer. A resolver now caches resolution failure itself for at least one second and no more than five minutes. That failure entry suppresses repeated upstream work. It is not a positive RRset or NXDOMAIN proof. A correct trace shows both objects: the expired answer eligible for fallback and the live failure state that explains why refresh is not being attempted on every request.

The stale response is a new projection

When returning expired data, the resolver must put a TTL greater than zero in the response; RFC 8767 recommends thirty seconds. That number is not the record's remaining authoritative lifetime. Ordinary lifetime is already over. It is a short downstream cache budget for the resolver's current projection of an older answer.

The distinction is easy to lose in logs. “TTL=30” could describe a genuinely fresh RRset near expiry or a stale answer assigned a new short response lifetime. Evidence therefore needs the original TTL, insertion time, ordinary expiry, stale age, response TTL and fallback reason. Without all six, an analyst cannot tell what the client was permitted to believe.

RFC 8914 supplies Extended DNS Error code 3 for a stale answer and code 19 for stale NXDOMAIN. The annotation is useful, especially across forwarding resolvers, but EDE does not change RCODE processing. It does not force a stub resolver to display a warning, refuse the answer or retry elsewhere. It is descriptive evidence, not a control plane and not proof that the old data remains safe.

Negative memory can deny a new future

Stale positive data keeps an old destination reachable. Stale NXDOMAIN keeps a name absent. These are not symmetric business risks.

Suppose an organization creates an emergency hostname during an attack. A recursive resolver holds an expired NXDOMAIN for that name and cannot refresh. Returning stale NXDOMAIN preserves the earlier denial precisely when the new record matters most. The client sees a clean negative answer rather than an obvious outage. EDE 19 can describe the condition, but many application paths will not surface it.

Signed denial increases the stakes. Stale NSEC or NSEC3 can continue covering a newly published name, DS or TLSA record. That may delay a delegation repair, a certificate policy change or a security upgrade. A stale answer is not always continuity; sometimes it is the continued execution of yesterday's absence.

That is why policies may need data-class exceptions. A broad resolver setting is simple, but certificate validation, delegation changes, emergency revocation and ordinary content addressing do not have identical costs of staleness. Any exception must be tied to observable record classes and tested implementations, not a vague promise that “critical DNS” will be fresh.

DNSSEC has its own clock

Storage lifetime and cryptographic validity are independent. RFC 4035 requires the validator's current time to fall within the RRSIG inception and expiration interval. Keeping a signed RRset in cache beyond TTL does not move the signature's expiration. RFC 8914 even reserves EDE code 7 for signature expiry.

A stale signed answer whose RRSIG is still within time may retain its validated security state under the resolver's rules. Once the signature interval ends, the same bytes cannot be authenticated merely because the maximum-stale timer continues. Operators need separate counters for stale-but-signature-valid, stale-with-expired-signature, validation failure served as failure, and any locally permitted insecure behavior. Collapsing them into “DNSSEC enabled” turns the most consequential boundary into a Boolean.

The threat is not only accidental. An attacker who can keep authorities unreachable may prolong an old address or denial. RFC 8767 notes the possibility of old hostnames pointing to addresses no longer controlled by the domain owner, with consequences for domain-validated certificates. Serve-stale does not create abandoned infrastructure, but it can enlarge the period in which old bindings remain actionable.

BIND and Unbound prove why option names are weak evidence

Current BIND documentation separates retention (stale-cache-enable) from returning stale answers (stale-answer-enable). It exposes response TTL, maximum stale age, refresh interval and runtime rndc serve-stale control. Its documented response TTL defaults to thirty seconds and maximum retention to one day when the feature is active. The current documentation also says the client-timeout option is off by default and, when enabled, supports only zero. That is an immediate-return behavior, not the full 1.8-second example merely because the configuration field exists.

Unbound uses the serve-expired-* family. Current documentation distinguishes immediate expired replies from RFC-style waiting before fallback, and separately configures expired age, reply TTL and reset behavior. Its documented defaults and version transitions are its own running-code history, not a DNS-wide promise.

A mixed resolver fleet therefore needs release-specific experiments. Send the same controlled query to each node. Make the authority fast, slow, unreachable, SERVFAIL, REFUSED and NXDOMAIN. Expire a positive record and a negative one. Keep an RRSIG valid, then let it expire. Record whether the node waited, what it returned, whether it attached EDE, whether it continued refresh and when it replaced or evicted the entry.

Sources