Summary

  • An RPKI manifest number is a local sequence signal for one CA's signed file inventory, not a global clock or a statement about what the network is doing. Its encoding has a maximum value, so defective issuance can leave later manifests unable to be greater than the number a relying party has already accepted.
  • RFC 9981 lets a different manifest filename establish a new number-comparison epoch. The relying party must alert its operator, and the issuer should not use filename changes as routine number management.
  • The epoch change resets only the remembered number baseline. Signature and certificate validation, a later thisUpdate, the file-and-hash inventory and exact correspondence between the signed-object URI and the RRDP or rsync location still have to pass.

The manifest that was newer but could never be greater

Imagine a certification authority whose publication system makes one catastrophic arithmetic mistake. Instead of issuing manifest number 7,914, it emits the largest number the field can encode: 2^159 - 1. The object is signed by the right key. Its time is current. Its file list and hashes accurately describe the publication point. Several relying parties accept it and remember the number.

The operator detects the defect, restores the counter and issues a corrected manifest. The replacement is later in real time, correctly signed and internally sound. It is also doomed under the old comparison rule. No legal value can be greater than the maximum already stored.

That is not merely an integer-overflow anecdote. It is a conflict between two kinds of evidence. Cryptographic evidence says the issuer made the new statement. Sequence evidence says the statement cannot follow the one already accepted. A validator that ignores sequence creates a replay path; a validator that refuses every later number can freeze the publication point forever. Both reactions protect something real.

RFC 9981 resolves the conflict by defining the smallest observable break in continuity that can start a new comparison history: a different manifest filename. It does not grant the issuer a general power to erase inconvenient state. The filename change is an epoch boundary, and all the other evidence survives it.

What the manifest accounts for

An RPKI manifest is a signed inventory issued by a CA for one publication point. It names the files the issuer intends to publish and binds each name to a hash. A relying party can use that inventory to notice that an object is missing, that an unexpected file has appeared in place of an expected one, or that the repository view is incomplete.

The manifest itself is an RPKI signed object. It follows the shared validation rules for RPKI signed objects and is carried by a one-time-use EE certificate. Its contents include manifestNumber, thisUpdate, nextUpdate, the file-hash algorithm and the filename/hash list. The CA certificate points to the manifest through the id-ad-rpkiManifest Subject Information Access description.

That evidence has a deliberately limited scope. A manifest can account for a ROA file without proving that a corresponding route is visible in BGP. It can account for a CRL without proving that every relying party fetched it. It can show what the issuer intended at a publication point without proving that a router consumed the resulting VRP. The manifest describes a repository claim at one layer; validation and routing consequences occur at later layers.

This distinction matters during recovery. If leaders describe the incident as “RPKI was wrong,” they will ask for a broad trust reset. The narrower diagnosis is that one issuer's sequence state became impossible under one filename. The correct repair can therefore remain narrow too.

A counter that was designed for continuity

RFC 9286 requires the issuer to increment the manifest number by one for each newly issued manifest. A relying party compares the observed value with the previously accepted value and expects the new number to be greater. The gap can reveal skipped publication state; a regression can reveal a stale or replayed manifest.

The number is not a global timestamp. Two CAs do not share one sequence. A large number from one issuer is not “newer” than a small number from another. Even within one CA, the sequence does not replace time. The relying party also evaluates thisUpdate and nextUpdate, certificate validity, revocation and the remaining signed-object rules.

The design assumed monotonic progress but not an unbounded integer. RFC 9981 makes the limit explicit: a positive INTEGER encoded in no more than 20 octets has a maximum of 2^159 - 1. Normal issuance will not approach it. At one manifest per second, exhaustion would take roughly 23,171,956,451,847,141,650,870 quintillion years. Capacity planning is not the issue.

Software is. A defective counter can add far more than one. An issuance loop can run without delay. Restored state can be corrupted. A value can be set to the maximum by mistake. The field is astronomically large and still operationally reachable in one bad assignment.

Once that value has been accepted, implementations face different failure clocks. One may retain the remembered state until an expiry condition and later recover. Another may continue rejecting lower numbers indefinitely. The same issuer can therefore look healthy to one relying party and permanently stale to another. The problem is no longer only publication; it is interoperability across remembered state.

The filename as an explicit epoch boundary

RFC 9981 says that when the manifest filename differs from the filename previously associated with the CA, a relying party must not reject the new manifest solely because its number is not greater. If the manifest passes the remaining checks, the relying party records the number under the new filename and begins a fresh comparison epoch.

The rule is precise because “the filename” is not a casual label chosen by a repository browser. It is the final path segment of the URI in the CA certificate's id-ad-rpkiManifest SIA. The CA must issue the certificate state that names the new location, publish the manifest at the corresponding location and keep the signed object's own URI binding consistent with the way it is retrieved.

This gives the transition three observable elements: the CA names a new manifest path, the repository serves the object there, and the relying party recognizes that the comparison key has changed. An operator can capture all three. No central coordinator needs to declare a universal reset, and no validator needs to guess that a lower number is “probably newer.”

The relying party must alert its operator when it sees the change. That alert is not decorative. A filename change may be legitimate recovery, a planned CA transition, an operator mistake or evidence of a more serious compromise. Human attention supplies the context that the signed bits cannot contain.

The issuer should not rotate the filename routinely. If every restart or release creates a new epoch, the monotonic number loses its ability to expose regressions, and attackers gain more opportunities to present an old but otherwise valid history as a fresh beginning. Exceptional authority has value only while it remains exceptional.

What the new epoch does not reset

The easiest operational mistake is to convert the RFC 9981 rule into “new filename means accept.” It does not. The filename answers only one question: which remembered manifestNumber baseline should be used for comparison?

The new manifest still needs a valid signature and a valid EE certificate chained to the CA. Certificate validity, revocation and resource-profile checks remain. The manifest's structure, hash algorithm and file list remain subject to RFC 9286. Missing files and hash mismatches retain their meaning.

Freshness also survives. RFC 9981 requires the new manifest's thisUpdate to be later than the thisUpdate of the previously accepted manifest. An attacker cannot take an old signed inventory, move it beneath a different-looking path and rely on the new epoch to defeat the time comparison. nextUpdate continues to bound the issuer's stated validity interval.

Location survives too. The one-time-use manifest EE certificate contains a signed-object SIA URI. That URI has to correspond exactly to the object's publication location: the RRDP publish URI when obtained through RRDP, or the rsync URI and path when obtained through rsync. A repository administrator cannot create a valid epoch merely by copying bytes to a new filename while leaving signed references pointed elsewhere.

Finally, the manifest does not rewrite the objects it lists. A new epoch can make a new inventory eligible for validation; it cannot change the resources in a certificate, the prefixes in a ROA, the state of a CRL or the content of a router feed. Those claims must validate on their own.

The multiple-SIA trap

The clean mental model—one CA, one old filename, one new filename—can fail when a CA certificate carries more than one manifest SIA access description. Different relying parties may prefer or support different locations. If the replacement certificate removes one old filename but leaves another, one validator observes a new epoch while another believes the old epoch continues.

The result is a split that can survive even though both validators process correctly according to the URI they selected. One accepts the low number under the new name. The other compares it with the poisoned maximum under the retained name and rejects it. Operators may then misdiagnose the disagreement as cache delay, RRDP inconsistency or product failure.

RFC 9981 closes the ambiguity: when relying on filename change for reset, none of the previous manifest filenames may remain in the new CA certificate. Recovery planning must therefore inventory every SIA value, every protocol path and every relying-party selection rule. Updating the “primary” repository endpoint is not enough.

The same discipline applies to publication. RRDP and rsync are delivery mechanisms, not alternative truths. Snapshot, delta and rsync views should resolve to objects whose signed URIs and bytes agree. A successful HTTP response proves transport. It does not prove that the fetched object is the one the certificate authorized.

Why a trust anchor is harder

A subordinate CA that poisons its manifest-number history can generally use the established CA key-roll procedure to create a new CA identity and a new manifest history under its parent. That operation still requires careful overlap, certificate publication and child continuity, but an authenticated superior remains available to bind the transition.

A trust anchor has no superior CA. Its public key and certificate location enter the relying party through bootstrap data, typically a TAL. Replacing that key by distributing a new TAL is slower and riskier: relying parties update at different times, old installations persist, and an abrupt cut can make the whole tree unreachable for validators that retain the old trust material.

RFC 9691 improves planned trust-anchor transition through TAK objects. A current trust anchor can advertise a successor public key and its certificate locations. Relying parties verify reciprocal current, predecessor and successor statements and observe an acceptance timer before moving. The design creates an in-band, testable bridge from the currently trusted key.

But TAK is not a magic side door around a damaged publication point. The relying party begins from the current trusted key, retrieves the CA certificate, validates the manifest and CRL, and then processes the TAK if present. A trust anchor that has not staged continuity may still face an acute bootstrap problem. The leadership lesson is to prepare key transition before it is needed and to keep manifest-number recovery as a separate, rehearsed procedure.

Recovery is proved by diverse validators

The issuer's repair is not complete when a new file is signed. It is complete when independently implemented relying parties fetch the exact object, accept the epoch transition under RFC 9981, validate the inventory and converge on the expected outputs.

A safe exercise begins with captured or synthetic repository material, never a production counter near the maximum. The test set should include a normal sequence, an anomalous jump, the maximum value, a lower number under the same filename, a lower number under a new filename, an older thisUpdate under the new filename, a mismatched signed-object URI and a certificate that retains one former SIA filename.

For each validator and version, the operator records the fetched URI, old and new filename, stored number, time comparison, alert, validation result, accepted object set and VRP delta. Differences are evidence. They show where implementation state or specification support diverges before a real issuer depends on the escape hatch.

Production recovery needs the same ledger. Preserve the former certificate, manifest, URI set, number, times and output fingerprints. Record the reason for the epoch change, the approvers, the new SIA values and the expected effects. Observe RRDP snapshot and delta behavior as well as rsync. Compare validator outputs and router-feed changes before removing old repository material.

The rollback boundary should be explicit. If the new manifest fails signature, time, URI or inventory checks, the answer is not another filename. Stop issuance, retain evidence and restore the last demonstrably valid publication state where protocol rules permit. Repeated resets turn a controlled exception into an unbounded ambiguity.

Sources