Summary

  • DNS greasing periodically sends unallocated extension values so servers and middleboxes cannot safely hardcode today's registry, but a reserved grease range can itself be specially allowed while all other unknown values remain blocked.
  • A greased failure proves intolerance somewhere on the tested path, not which component caused it; retries can restore service while erasing the first failure unless telemetry preserves both events.

The test was green. A DNS resolver sent a value from the range labelled for greasing, the answer came back, and the path was declared extensible. Months later a real extension received a new code point outside that range. The same path rejected it immediately.

Nothing in those results is contradictory. The middlebox had learned the test.

This is the paradox inside revision 03 of Greasing Protocol Extension Points in the DNS. Greasing tries to prevent protocol ossification by making unknown values ordinary before they acquire meaning. If an intermediary cannot assume that every unfamiliar code is an error, future extensions have room to arrive. But if the test values become predictable, the intermediary can hardcode one more exception and preserve the underlying prohibition.

Revision 03 is an active DNSOP Working Group Internet-Draft dated 6 July 2026 and expiring 7 January 2027. It is intended Informational, not a final RFC, implementation report or measurement of deployed DNS. The text openly leaves important work unresolved: whether to reserve code-point ranges, which registries could spare them, detailed behavior, fallback, settings and telemetry. Those gaps are not editorial trivia. They define who will own the evidence when a preventive test affects real traffic.

Unknown values keep an extension point alive

Protocols often contain fields designed for future values. A registry assigns some values while others remain unallocated. Conforming software is expected to tolerate an unknown value in particular locations—often by ignoring it—so a later allocation can be deployed without first replacing every intermediary.

That promise decays when implementations see only the assigned set. Servers validate against a frozen list. Firewalls discard packets containing anything unfamiliar. Proxies parse a field more narrowly than the protocol allows. Their behavior may remain invisible for years because normal traffic never tests the unused space. When a real extension finally arrives, the path has already hardened around its absence.

Greasing reverses the timing. A resolver periodically uses an unallocated value while it still has no operational semantics. Correctly tolerant receivers continue. Intolerant paths fail early enough to be discovered and repaired before the value carries an indispensable feature.

The draft identifies several DNS surfaces: the remaining DNS header flag position, Opcode, EDNS Version, EDNS header flags, Class, Resource Record Type and EDNS Option Code. Their available spaces are radically different. A one-bit flag collection cannot be governed like a 16-bit registry with tens of thousands of unallocated values. RR Type space also has internal data, meta and query-type structure. One policy cannot be projected across all of them.

The draft excludes RCODE and Extended RCODE from the candidate list. Those values are meaningful only when set by a responder and directly change how the querier interprets status. Injecting an arbitrary error code would not simply test tolerance; it would alter the outcome. Greasing is safe only where “unknown” already has a bounded, conforming treatment.

Reserved ranges solve one problem by creating another

Randomly selecting from currently unallocated space has an important advantage. The receiver cannot easily know whether a value is a test or the beginning of a real deployment. It must implement the general unknown-value rule.

Random selection also carries a custody risk. A value unallocated when software ships may be allocated later. Old software can continue emitting it as grease, or a receiver trained to ignore it can mishandle its new semantics. The registry has changed, but the deployed binary has not.

A reserved grease range quarantines test values. Future allocations cannot collide with it, and operators can recognize tests in telemetry. Yet recognition is also the weakness. A middlebox can permit the reserved range and reject all other unknown values. The grease test passes while the extension point remains closed.

The reserved range has become a ceremonial doorway. It proves that the path tolerates the ceremony, not that it tolerates change.

This is why revision 03 does not announce a settled answer. It records arguments for reservation—diagnosis and collision avoidance—and arguments against it—self-ossification, deliberate special-casing and the scarcity or internal structure of some registries. The Working Group has not reached consensus. The reserved-values section remains a placeholder, and no IANA action can be inferred.

An honest deployment report would therefore avoid a single “grease compliant” status. It would name the extension point, selection policy, tested values, registry snapshot and whether the values came from a known reserved set. Passing one range is not evidence about another.

A failure belongs to a path, not automatically to a server

Suppose a greased query times out. The authoritative server might reject the unknown value. So might a firewall, DNS proxy, load balancer or transparent middlebox. An anycast path could deliver the test to a different node from the control query. A transport difference or cache state might change the comparison.

The observed fact is narrower: this query, carrying this value, did not produce the expected interoperable result along this path at this time. Naming a broken authoritative implementation requires more evidence.

Useful telemetry binds the destination, query, extension point, grease value, transport, timestamp, response or timeout and comparison attempt. If operators aggregate results, they also need the sampling frame and denominator. Without those receipts, a failure counter becomes accusation without location.

The distinction matters for repair incentives. A zone operator cannot fix a resolver-side parser. A resolver vendor cannot repair an enterprise firewall it cannot identify. Telemetry that blames the visible endpoint may create pressure while sending it to the wrong principal.

Fallback preserves availability and can erase the crime scene

Revision 03 says resolvers should generally retry failed queries without the unallocated extension. The exception is a test that constructs a different query, such as asking for a new RR type, where removing the value no longer means repeating the same operation.

A plain retry is sensible service protection. The user receives an answer even if the path is intolerant. But if monitoring stores only the final success and total latency, the fallback has destroyed the reason greasing exists. The operator sees a slightly slower resolution, not a path that will block a future feature.

The evidence model needs two linked events: the greased attempt failed under a defined condition; the ungreased attempt succeeded under the changed condition. The second is not a correction of the first. It is a control observation.

Parallel queries can avoid user-visible retry delay. The resolver sends one normal query for service and one greased query for telemetry. Only a small sampled fraction of traffic needs duplication. This can strengthen comparison, but it also increases volume and does not guarantee both queries hit the same anycast instance, cache state or network path. A pair needs a binding stronger than “same name around the same time.”

Automatic fallback has a security edge too. If a future extension carries a security property, an attacker who can induce its failure may train the resolver to retry without it. The availability mechanism becomes a silent downgrade surface. Greasing is intended to expose such intolerance early so operators can eventually remove unsafe fallback, not to normalize fallback forever.

Sampling is not an ecosystem census

The draft suggests sampling because high-volume resolvers need only a small fraction of queries to obtain a rough picture. “Perhaps 1 in 1000” appears as an illustrative question, not a mandated rate.

One in a thousand popular queries can still be a biased sample. It may overrepresent large authoritative platforms, common names, certain geographies, a resolver's own customer mix and peak-hour paths. Rare legacy systems—the places most likely to resist change—may appear too infrequently for a stable estimate.

Community aggregation could broaden visibility, but pooled counts do not create a neutral denominator. Operators use different selection policies, transports, retry logic, anycast footprints and reporting thresholds. A global percentage without those dimensions borrows authority from scale while hiding incomparable inputs.

The right output may be a map of tested populations and failure conditions rather than one headline number. Provenance lets other operators decide whether a result applies to their path.

DNS Error Reporting could notify authoritative operators about observed deficiencies. Such reports remain claims from a reporting resolver and require their own delivery and anti-spoofing controls. A report is a lead for investigation, not a remote verdict.

The responder cannot see the downstream result

Most of revision 03 focuses on grease initiated by the resolver or querier. A responder could also return an unknown flag, option, EDNS version or RR type. That direction is more difficult because the server may never learn what happened next.

The querier might accept the answer normally. It might reject it and retry another authoritative server. It might fail without retry. The downstream application might abandon the operation. Silence at the first server is compatible with all four outcomes.

Targeted responder experiments can still be useful, as earlier delegation measurements have shown. But they need an observation channel, controlled clients or user reporting. Emission is not acceptance. The draft therefore says responder-initiated greasing should be disabled by default.

That asymmetry is an important governance constraint. The actor injecting the test does not always control or observe the affected result. A safe experiment needs an impact budget and a way to close the evidence loop.

Exit is part of the experiment

The draft offers an end-of-test date as one way to reduce collision risk when random unallocated values are used. After that date, a particular grease operation retires before later code assignments create conflict.

A date in source code is not retirement. Deployed software must receive the update or enforce the deadline locally. Long-lived appliances can continue emitting a value after it acquires real meaning. Operators need to know the current registry state, software release state and active configuration at packet time.

Predictable non-random selection has another cost: fingerprinting. Observers can correlate a distinctive sequence with a resolver product or version. Good variability reduces that signal, but any measurement design has privacy and volume consequences of its own.

Greasing succeeds when future change becomes ordinary, not when a reserved test produces green dashboards. The enduring evidence is not that one value passed. It is that unknown-value behavior remains general, failures stay visible through fallback, sampling keeps its provenance, and the experiment can truly stop before the test value becomes somebody else's production semantics.

Sources