Summary
draft-ietf-dnsop-dnssec-keyrestore-02addresses a narrow catastrophe: a private signing key becomes inoperable, no usable backup exists, but a complete pre-signed zone and still-valid signatures remain.- Safe recovery follows different clocks for signature expiry, child DNSKEY propagation, parent DS publication, secondary loading and cache retirement. A signer can start several transitions but cannot complete all of them alone.
- Operators should keep a recovery-clock receipt that records the authority, evidence and earliest safe time for each transition. This is Daniel Kade’s governance recommendation, not an IETF requirement.
The alarming screen is not always the public outage. A hardware security module fails, an operator loses access to a key, or a disaster makes the private material unusable. The authoritative servers continue to return a signed zone. Validating resolvers still find a DNSKEY, an RRSIG and, where relevant, a parent DS that fit together. Users may notice nothing.
This quiet interval creates two opposite risks. Panic can make a valid zone bogus before the old signatures expire. Comfort can waste the interval until the signatures that conceal the failure are about to run out. Good recovery therefore begins by naming the state accurately: the service has retained validation continuity but lost its capacity to sign new state.
Revision 02 of DNSSEC Key Restore was published on 10 August 2026. It is an active DNSOP working-group Internet-Draft, intended as Informational in the rendered document, and remains work in progress. It is not an RFC, a BCP or a deployment commitment by the authors’ employer, RIPE NCC. The draft concerns pre-signed zones. Online signing and the root zone are expressly outside its scope.
Its title is deliberately modest. The procedure does not restore a vanished private key. It restores the signing function by introducing replacement key material while preserving the public records and old signatures on which validation still depends.
A functioning answer can hide a frozen zone
The draft defines an inoperable private key as a private part that can no longer produce signatures. Hardware failure, natural disaster, operator error and malicious action are possible causes. A compromised key is a different condition if it can still sign. That distinction matters because compromise raises an adversarial-use problem; inoperability raises a continuity problem. Treating them as synonyms can produce the wrong response.
At the instant of failure, the zone may still be correctly signed and served. RRSIG lifetimes often leave days before validation fails. “Often” is not a service guarantee. The remaining interval comes from the actual inception and expiration times of signatures across the zone, the records still published, and the combinations that validators may have cached.
During that interval, the zone may be unable to change safely. A new address, revoked service, mail-routing fix or incident-response record cannot simply be inserted if the relevant old key cannot sign the changed RRset. The absence of a visible resolution failure therefore says nothing about change readiness. A stable answer can be evidence of preserved old state, not evidence that normal operations continue.
The draft’s central precaution follows. Signing software must not discard old DNSKEYs merely because their private halves are absent. It should retain old RRSIGs, and an operator may have to add them back manually if the signer would remove them. Old public material is no longer operationally productive, but it remains validation-critical until every dependent cache horizon is over.
The recovery path changes with the key role
DNSSEC validation treats keys cryptographically alike, but recovery does not. The authority path differs according to whether the unusable key is a Zone Signing Key, a Key Signing Key or a Combined Signing Key.
If the ZSK fails while the KSK remains operable, the operator can add a new ZSK to the DNSKEY RRset and use the working KSK to sign that set. The new key does not become ready merely because it appears on one authoritative server. The draft’s publication interval combines authoritative propagation delay and the DNSKEY TTL: Ipub = Dprp + TTLkey. Only after that horizon is the replacement assumed to be available wherever a cached DNSKEY set may matter.
The zone can then be signed with the new ZSK. Retirement is another clock. The old ZSK remains until new signing has completed, changes have propagated, and signatures made with the old key can no longer remain in caches: Iret = Dsgn + Dprp + TTLsig. The key failure, key publication, signing activation and safe removal are four different events.
If the KSK fails while the ZSK remains operable, the child cannot solve the problem inside its own zone. Trust in the replacement KSK must be connected through a DS record in the parent. The draft uses the Double-DS method: submit a new DS, wait for parent registration and publication, allow the new DS to propagate and age into caches, then activate the new KSK in the child.
The parent publication time follows the submission time plus a registration delay, Tpub = Tsbm + Dreg. Readiness then adds parent propagation and the DS TTL, IpubP = DprpP + TTLds. RFC 7583 already makes the governance point plainly: KSK timing involving a parent is not wholly controlled by the zone manager. An API receipt saying “accepted” cannot shorten a parent’s publication or a resolver’s cache.
If a Combined Signing Key fails, both surfaces meet. The replacement needs the parent-side Double-DS path, while the zone also lacks its ordinary signing capability. Old CSK records and signatures must remain until the new DS, new key and newly signed zone have moved through their respective horizons. The retirement interval includes signing delay, child propagation and the larger of the DNSKEY and RRSIG TTLs. The apparent simplicity of one key becomes a broader recovery dependency.
Failure turns automation back into ceremony
Healthy DNSSEC operations can automate parent DS maintenance through CDS or CDNSKEY. In the failure state, the records that would ask for a new DS may themselves be impossible to authenticate with the inoperable ZSK or CSK. Revision 02 therefore sends the affected recovery path back through a manual parent update.
Manual does not mean informal. Someone must identify the authorized child, authenticate the request, verify the replacement material, approve an exception to the ordinary automation and distinguish receipt from publication. That chain may cross a DNS host, registrar, registry and parent-zone operator. The cryptographic failure exposes the institutional path that routine automation had hidden.
Secondary service creates a second manual seam. Introducing a new DNSKEY while keeping the SOA unchanged avoids changing a record that the old key can no longer sign. But an unchanged SOA may give secondaries relying on IXFR or AXFR no ordinary signal to fetch the altered content. They may need to be forced to load it. One successful primary response is not evidence that every authoritative instance carries the same recovery state.
ZONEMD adds another consequence. Changing the DNSKEY RRset changes the zone contents and invalidates the existing digest. A new digest must be generated and signed. The draft’s Knot DNS implementation note says manual generation of that digest is not supported in the described procedure. That is useful evidence of a boundary, not proof that the method is complete across products.
At IETF 126, participants asked for closer comparison between TTLs and signature expiration, and for implementation detail. The presenter reported that the approach worked with Knot DNS; the chair requested an implementation section. The resulting draft evidence is meaningful but narrow. It should not be inflated into a multi-vendor interoperability claim.
A recovery-clock receipt
An incident ticket usually records a failure time and a resolution time. That is too coarse for this recovery. Between those endpoints sit transitions owned by different actors, each with its own evidence and earliest safe time.
A recovery-clock receipt should begin with the incident classification, affected key role and algorithm. It should calculate the earliest RRSIG expiration that bounds continued validation, rather than rely on a usual signing lifetime. It should preserve proof that the old DNSKEY and signatures remain served. It should record the new key’s publication across authoritative instances and the propagation and TTL inputs used to decide when the key is ready.
For KSK or CSK recovery, the receipt should separate parent submission, authentication, acceptance, observed publication and cache-ready time. Those are not synonyms. For secondaries, it should name every forced-load target and the observed content or serial state, while acknowledging that an unchanged SOA complicates the normal comparison. For ZONEMD, it should record whether a new digest was generated, why it was temporarily unavailable, and who approved that state.
Finally, the receipt should identify when new signing completed, show a bounded sample of newly signed RRsets, calculate the earliest retirement time for old DNSKEY, DS and RRSIG material, and record actual removal separately. Each manual exception needs a decision owner and a correction route.
The purpose is not to publish secrets. Private keys, HSM configuration and personal incident data do not belong in a public record. Hashes of artifacts, key tags, timestamps, policy versions, observed RRsets and responsible organizational roles can establish a useful chain without exposing the signing environment.
Measurement must admit what it cannot see
No operator can inspect every recursive cache. TTL-based timing therefore uses conservative horizons, not universal observation. Independent queries from multiple vantage points can show that sampled authoritative and recursive paths have reached the expected state. They cannot prove that no stale combination exists anywhere.
That limitation is a reason to preserve formulas and inputs. If a reviewer knows Dprp, TTLkey, TTLds, TTLsig and the signing delay assumed by the operator, the reviewer can test whether the transition respected the chosen safety margins. A green dashboard without those inputs turns a calculation into an assertion.
The state vocabulary matters equally. A request can be submitted but not accepted; accepted but not published; published but not ready; ready but not active; active while old material is not yet safe to remove. Compressing these into “recovered” moves the conclusion ahead of the evidence.
What the draft proves—and what it does not
Revision 02 provides a technically grounded way to recover signing capability without intentionally making a pre-signed zone insecure or bogus. It demonstrates why retaining old public material and respecting cache horizons are part of recovery, not housekeeping after it. Its formulas make several dependencies explicit.
It does not provide a universal recovery time, audit format, parent-service promise or resolver census. It does not apply to every signer architecture. It does not say that every manual step is safe merely because automation failed. And it does not replace the distinct response required when a key is compromised.
The governance conclusion follows from those limits. Old signatures buy a temporary continuity budget. They do not give one party control over the parent, secondaries or caches. A recovery process becomes accountable when each clock is tied to the authority that can move it and the evidence that proves it moved.
Sources
- DNSSEC Key Restore — revision 02
- Datatracker record for DNSSEC Key Restore
- DNSSEC Key Restore document history
- IETF 126 DNSOP minutes
- DNSOP active documents
- DNSOP charter
- RFC 9364: DNS Security Extensions
- RFC 7583: DNSSEC Key Rollover Timing Considerations
- RFC 8078: Managing DS Records from the Parent via CDS/CDNSKEY
- RFC 8976: Message Digest for DNS Zones
- RFC 9499: DNS Terminology
Member Briefing
Deeper Profile Context
Sign in with the right membership level to unlock the full briefing and source notes.
Only for Strategic Circle
Strategic Circle
Open to all readers. Unlock profile briefings after joining and signing in.
Join Strategic CircleOnly for Leadership Alliance
Leadership Alliance
For qualified IP-asset owners and management; sign in to unlock alliance briefings.
Join Leadership Alliance
