Summary

  • draft-kaizer-dnsop-ml-dsa-mtl-dnssec-02 lets many DNS RRsets share an ML-DSA-signed Merkle ladder, replacing repeated full signatures with authentication paths plus at least one signed ladder.
  • The proof can validate in one response, but it does not prove that authoritative signers preserved the node-set lineage needed for updates, transfer, failover or recovery; the draft explicitly leaves those operating rules to later work.

Revision 02 appeared on 28 September 2026. Its header now targets the Standards Track, and it proposes an IANA registry for MTL types. It is still an individual Internet-Draft: not a DNSOP Working Group adoption, not IETF consensus and not an RFC. The requested algorithm number is still TBD. Four named open-source projects show that testing code exists, but the draft itself says those contributor-supplied listings are unverified and imply no IETF endorsement.

The problem is real. ML-DSA gives DNSSEC a post-quantum signature standardized by NIST, but its signatures are much larger than the ECDSA signatures common today. Repeating one full ML-DSA signature for every RRset expands zones, caches and responses. Merkle Tree Ladders try to pay that cost less often.

The draft places an ML-DSA public key in DNSKEY. Instead of signing each RRset independently with ML-DSA, the signer hashes each DNSSEC message into a leaf. Leaves receive consecutive indexes beginning at zero. They accumulate in a node set identified by a 32-octet series identifier, or SID. A changing set of rungs authenticates those leaves. The expensive ML-DSA operation signs the ladder; an individual RRSIG carries its leaf index, randomizer, authentication path and the signed ladder needed for a full response.

A resolver can therefore verify one response without asking for a second object. It validates the signed ladder against the DNSKEY, reconstructs the leaf hash from the RRset, follows sibling hashes to a compatible rung and then continues the ordinary DNSSEC chain. One full signature can cover every RRset represented by that ladder. The economy is genuine.

So is the state. The ladder is not a generic certificate that floats above the zone. It is a view of one evolving node set. The SID identifies that series. The leaf index says where a message entered it. The randomizer contributes to the leaf hash. The current rungs depend on the number and order of appended leaves. When the node set grows, an older message may need a new authentication path relative to the new ladder.

That does not mean an implementation must preserve one proprietary database forever. State may be replicated, reconstructed from sufficient material or retired by beginning a new series. The current draft simply does not define which method is interoperable or safe. It says later documents may cover zone signing, composition, updates, transfer, name-server processing, resolver processing and caching. A cryptographic format has been specified before its continuity contract.

Consider a hidden-primary cluster. Signer A appends an RRset at leaf 48, signs a new ladder and distributes the zone. Signer B has the same private key but a backup taken when the counter was 47. If B becomes active, several outcomes are possible: it might reuse an index with different material, rebuild a different ladder, abandon the old SID, or refuse to sign. Which outcome is correct depends on state and recovery rules that are not in revision 02. A resolver validating yesterday's proof cannot answer whether tomorrow's signer will extend the same history coherently.

Key custody is therefore necessary but insufficient. The draft requires different SIDs for separate MTL instantiations, explicitly including KSK and ZSK node sets. An operator must know which series belongs to which key role, which signer owns the next index, how a ladder generation is made current, and how stale state is rejected. Copying only the private key and unsigned zone data may not copy the decision history embodied in the node set.

Batch signing makes that history economically useful. Adding several RRsets before producing a new signed ladder reduces average ML-DSA operations and can lower HSM load. It also creates a timing decision: wait too long and changed data lacks a fresh usable proof; sign too often and the amortization shrinks. The draft correctly leaves batch size to the zone operator. A service objective must nevertheless say how long a new or changed RRset may wait and which ladder becomes authoritative across the fleet.

Online signing exposes the same trade. The draft says an implementation could create a new MTL node set for each query response. That avoids sharing a long-lived series across requests but may confine the amortization benefit to one response. “Supports online signing” is thus not a performance result. It is a design space whose CPU, HSM, latency and recovery properties still need measurement.

The wire path is not small either. Revision 01 removed the EDNS(0) option that had supported an optimisation and made the full MTL type the only response type for every RRset. Revision 02 states that these full responses always exceed DNS over UDP. It expects additional TCP traffic and recommends TCP-first queries for clients that want to avoid truncation followed by retry.

TCP is a transport answer, not a continuity answer. A successful connection does not prove that every authoritative node has the same ladder state. Receipt of a full response does not prove that a resolver implements the proposed algorithm. A valid Merkle path does not prove that the DS/DNSKEY chain reaches a configured trust anchor. DNSSEC validation does not prove that the application used the answer or that the service was available.

The already-published SigTag proposal addresses a different step. It would let a client say which signed ladders it believes it has cached so a server can sometimes omit retransmitting one. That article owns cache claims, full-versus-condensed selection, privacy and fallback. The base draft in this article owns the state beneath that optimisation: before a client can reuse a ladder, authoritative operators must be able to create, preserve, identify and recover a coherent one.

This distinction follows running-code primacy. A document can define deterministic proof verification, and test programs can exercise it. Neither creates an operating fact. Adoption begins when independent implementations survive updates, role separation, cluster failover, zone transfer, rollback and realistic TCP load. A proposed code point records a format; it does not certify that history is portable.

The right evidence is a signer-state receipt. For each generated ladder, record the key role, SID, index range, rung set, zone serial, signer identity, software version, creation time and predecessor. Test that a second signer can continue or deliberately replace the series without reuse or ambiguity. Preserve enough material to recompute an old proof, or explicitly document why a new series invalidates that expectation. A recovery exercise should finish with resolver validation from a real trust anchor, not merely a process exit code.

Sources