Summary
- On 13 November 2025, ARIN administratively blocked all access to its Hosted and Repository Publication Service RPKI repositories from 13:30 to 14:30 EST. It reported restored access at 14:30, full redundancy at 14:35 and a return to pre-test levels at 14:50.
- The exercise tested recovery from a repository-access outage. It did not, in the public record, identify the external dependency classes, common failure domains, observer set or circular dependencies raised in ACSP Suggestion 2025.7.
- ARIN can close that gap without publishing attack-ready topology: a versioned service-to-failure-domain statement can record the protected service plane, dependency class, exercise boundary, observed milestones, residual assumptions and change-notification trigger.
Four dates make the evidence more interesting than a routine continuity announcement.
On 23 October 2025, ARIN received a customer report about its Hosted RPKI service after deployment of ROA Support for Transfers. Its public incident report, issued on 27 October, says the problem affected one customer under a specific ROA configuration condition. ARIN paused in-process transfers, identified a code defect, deployed a fix and added validation.
On 28 October, Ramakant Pandrangi submitted ACSP Suggestion 2025.7. It did not ask merely whether ARIN had backups. It asked for service-by-service disclosure of dependencies, including CDN, DNS, IP addresses, BGP routes, providers and other infrastructure; notice before material operational or architectural changes; and high-level continuity objectives such as RTO and RPO.
ARIN answered on 12 November. It said it was developing a plan to document key service dependencies, clarify change management and outline resilience and continuity objectives, with community feedback and validation to follow. The next day, ARIN ran a production failover exercise.
That sequence offers a rare opportunity to say exactly what public evidence proves—and where it stops.
The exercise produced five useful timestamps
ARIN’s technical-list notice says access to sites serving the Hosted and RPS repositories began to be restricted at 13:00 EST on 13 November. At 13:30, all repository access was administratively blocked to simulate a full outage. Access returned at 14:30. ARIN confirmed full redundancy at 14:35 and said the system reached pre-test levels at 14:50.
These are not marketing adjectives. They are observable milestones attached to a named service and a defined intervention. The notice also explains why ARIN chose production: a test outside production would not accurately represent performance and preparedness for repository failure conditions.
That deserves credit. Many resilience statements offer only “high availability,” without a dated exercise or a recovery sequence. ARIN supplied the sequence.
The result remains bounded. “All access was administratively blocked” identifies the condition presented to the service. It does not tell a reader whether the blocked sites shared a DNS provider, CDN control plane, transit path, facility, power domain or identity service. It does not state what failed underneath the intervention, because the intervention itself was administrative. Nor does the notice identify the validators or relying parties used to confirm recovery, or measure routing effects.
None of those omissions invalidates the exercise. They define its claim.
A recovery path is not a dependency map
Suggestion 2025.7 is concerned with systemic and circular dependencies. Those are questions about whether two apparently separate paths still require the same upstream condition. Two repository sites may be redundant at the application level while sharing a provider, a naming layer or an administrative control. Conversely, a service may depend on an external platform and still be resilient because the dependence is diversified, cached, replaceable or deliberately isolated.
A provider name alone would not answer the question. A list reading “DNS, CDN, transit” would be too broad. What matters is the invariant each service must preserve when a dependency class is unavailable.
For an RPKI repository, that invariant may concern continued retrieval of already published material, the age of the last accepted publication, integrity of repository state, or the ability to reconcile queued updates after restoration. The correct invariant depends on the service plane. A public reader, a resource holder changing a ROA, a delegated CA sending material to RPS and an ARIN operator restoring the service are not performing the same action.
ARIN’s own July 2026 maintenance notice illustrates the distinction. ARIN Online, RESTful Provisioning, RPKI Up/Down and RPS were unavailable, while the RPKI repository remained operational but published no updates. The earlier Theo March analysis of that notice concerned two clocks—availability and freshness. The present question is different: which failure domains can stop each plane, which combinations were exercised, and which remain assumptions?
Responsibility is deliberately split
ARIN’s deployment documentation makes the dependency question operationally material without making dependence itself suspect.
More than 95 percent of ARIN RPKI deployments use Hosted RPKI, according to its current page. In that model, ARIN operates the certificate authority and publishes the high-availability repository. A delegated participant can instead operate its own CA and publication server. RPS divides the duties again: the resource holder retains its CA and private-key control while ARIN operates the repository.
ARIN’s RPS page explains why. A consolidated repository can sit outside the administrative scope of the organization issuing an RPKI object. Some issuers want cryptographic control but do not want to maintain a 24/7 repository. Outsourcing that function is a rational separation of work.
The split creates two different continuity questions. Can the issuer continue to create valid publication material? Can the repository continue to make accepted material available? A test of repository access says something important about the second. It does not automatically test the issuer’s CA, the publication protocol, the account path, the notification path or an external service shared by several of them.
The distinction is particularly important in RPS. Cryptographic authority and publication availability belong to different administrative actors by design. Treating one as proof of the other would erase the very separation the service offers.
Availability is evidence, not architecture
ARIN’s 2025 annual report says its RRDP and Rsync repositories each achieved 99.9 percent availability. The figure is useful. It gives readers a reported performance measure over a defined reporting year, and it can be compared with later years.
But 99.9 percent does not show whether two paths share a failure domain. It does not disclose an RTO, an RPO or the age of data during a read-only interval. It does not say which exercise produced which part of the result. Availability can be excellent even when the dependency model is unpublished; a sound dependency model can exist even after a measured interruption.
The October incident report makes another distinction. That event followed a software deployment and was attributed to a bounded code defect, not to an external infrastructure provider. One customer was affected, according to ARIN, and the service otherwise continued to operate. It is useful evidence of change control and incident response. It should not be recycled as evidence that a CDN, DNS or repository failover path failed.
This is the discipline the public record needs: one claim per piece of evidence.
Publish the failure-domain statement, not the wiring diagram
There is a serious counterargument to dependency transparency. Detailed supplier names, facility locations, routing designs, control endpoints and failover triggers can reduce security. A public document should not become a reconnaissance guide. Supplier contracts can also contain confidentiality provisions, and a precise internal map can go stale faster than it can be safely reviewed.
That does not force a choice between silence and exposure. ARIN could publish a compact, versioned statement for each critical service and plane. It would contain:
- the public function and the readers or writers who depend on it;
- the dependency classes and common-mode assumptions, without naming sensitive vendors or locations;
- the invariant expected during loss of each class;
- the scope and injection boundary of the last exercise;
- the observed restriction, restoration, redundancy and baseline timestamps;
- the residual failure modes not covered by that exercise;
- the applicable availability, RTO and RPO objective, where adopted;
- the material-change threshold and member-notification route; and
- a version, review date and correction history.
Such a statement would make the November exercise more valuable, not less. Readers could see that the repository-access condition was tested, when it was tested and which wider claims the result did not make. Later exercises could cover a different dependency class without pretending that every scenario must be tested at once.
The document should also expire. A dependency assertion without a review date can survive an architecture change and become more misleading than no assertion at all. Suggestion 2025.7 expressly connects transparency with change management. The connection is sound: the evidence record must change when the service boundary changes.
The open status should remain a status, not a verdict
At the evidence cutoff, Suggestion 2025.7 still appears as Open. ARIN’s response says the work will be presented to the community when the plan is finalized. That supports a precise statement about the public process. It does not prove that internal planning has stalled, that no dependency inventory exists or that a resilience weakness is being concealed.
The next useful public milestone is therefore not a promise of perfection. It is a first bounded map: service planes, failure-domain classes, evidence dates and known limits. Community review can then challenge whether an invariant is meaningful without demanding sensitive topology.
ARIN’s failover test already shows the right instinct. It named the service, intervened in production and published the recovery sequence. The dependency question remains open because the exercise and the question live at different levels. Joining them carefully would turn a good operational result into durable institutional evidence.
Sources
Member Briefing
Deeper Profile Context
Sign in with the right membership level to unlock the full briefing and source notes.
Only for Strategic Circle
Strategic Circle
Open to all readers. Unlock profile briefings after joining and signing in.
Join Strategic CircleOnly for Leadership Alliance
Leadership Alliance
For qualified IP-asset owners and management; sign in to unlock alliance briefings.
Join Leadership Alliance
