Summary
- JANOG's temporary RPKI Routing WG ran bounded certificate, ROA, cache and router exercises in 2013, then publicly examined failure scenarios. It did not allocate resources, issue production certificates, set national routing policy or prove that an all-
Invalidevent had occurred. - RFC 6810 already specified multiple-cache and retained-data behaviour before the JANOG32 discussion. Later Japanese deployments, guidance and incidents converged on the same operational concerns without proving that the working group caused them.
- The 2022 JPNIC incident invalidated nearly all JPNIC-issued ROAs as repository entities; affected routes were observed as
NotFound, not route-stateInvalid. No route, traffic, relying-party or end-user-outage denominator was published. - Safe deployment requires layered monitoring, diverse failure domains, staged and reversible policy, local exceptions, and denominators that show what actually changed. Those safeguards strengthen the case for origin validation rather than justify indefinite delay.
The disk that changed the meaning of a route
On 26 January 2022, a disk filled up in part of JPNIC's RPKI repository system. The event was mundane in its mechanism and unsettling in its reach. Fresh certificate revocation lists and manifests were not published. As those entities aged, nearly all route origin authorisations issued by JPNIC became invalid at the repository-entity layer. A user noticed the problem. JPNIC manually recovered the system on 2 February and published an incident notice.
The crucial downstream result in that notice was NotFound. Routes affected by the loss of usable authorisation data were observed with that route-origin validation state. This is not the same result as a BGP announcement being Invalid. Nor does either label, by itself, show that a router rejected a path or that an end user lost service. The notice gave no count of affected routes, relying-party installations, traffic flows or end-user outages. It did not attribute the incident to JANOG. Its evidence was both serious and bounded: a repository publication failure had made nearly all JPNIC-issued ROAs unusable, and the observed downstream consequence was the absence of covering validated payloads for affected routes.
That distinction can feel pedantic until policy acts on it. An invalid repository entity belongs to the certification and publication layer. A validator's incomplete set of validated ROA payloads, or VRPs, belongs to the relying-party layer. NotFound and Invalid are route-validation states produced by comparing a BGP announcement with the VRPs available to the router. Rejecting, preferring, tagging or accepting the route is a local policy action. Lost connectivity is an observed outcome at still another layer. A sentence that collapses those steps can manufacture an outage, an attacker or a responsible institution that the evidence never established.
Nine years before the disk filled, a temporary working group in the Japan Network Operators' Group had already been asking what happens when the signal intended to improve routing trust becomes unreliable. It did not foresee this particular incident, and it did not invent the protocol mechanisms that could contain one. Its contribution was smaller and more useful: it put tools into operators' hands, found rough edges, and left a public record of the questions that appear when validation data is allowed to influence routing.
The safety question inside RPKI is therefore not whether trust data should matter. Origin authorisation is useful precisely because it can distinguish an announcement consistent with a resource holder's published intent from one that is not. The question is how to make that signal operationally consequential without pretending it is complete, infallible or controlled by one actor. Because each layer belongs to a different actor, resilience has to be assembled across several decisions. Failure also has to be reported precisely enough for operators to know which decision should change.
Six months, not a national authority
The JANOG RPKI Routing WG began on 22 January 2013 and ended on 31 July that year. Those dates matter. They define a temporary forum, not a permanent institution with authority over Japanese routing. The working group convened experiments, hands-on sessions, a tutorial and discussion among entities. The public record reviewed for this article does not establish a separate legal personality, contracting power or financial control for it.
JANOG, through the working group and meeting archive, created a place for operators to learn and compare experience; its documented role stopped at convening. JPNIC was the Japanese Internet registry and later operated the trial, repository and guidance surfaces described here. APNIC and other RIRs performed certification and repository roles within their service regions. Software projects and vendors implemented validators, cache servers, the RPKI-to-Router protocol and router behaviour. Each network operator chose its own cache topology, router configuration, exceptions and routing action.
Resource allocation, certification, regulation, standards-making and production operation belonged to those respective actors, not to JANOG.
Collapsing those roles would change the story. “JANOG deployed RPKI in Japan” would turn a bounded community exercise into a national operational act; treating JPNIC's later trial as a JANOG service would erase the registry's function; and treating a displayed validation state as a filtering decision would assign the operator's choice to the router. The history is distributed because the system itself is distributed.
The working group's first practical surface consisted of two hackathons in early 2013. Its own revised activity report, entity decks and a later JPNIC newsletter account describe work with RPKI tools, caches, certificate and ROA operations, and a path toward observing validation information on a BGP router. Entities used designated resources and a mock or experimental environment. The first round encountered environment and software defects. Developers made repairs, and the second round achieved better cache completion. Later hands-on sessions in April and May simplified the environment, including prebuilt virtual machines, so entities could issue certificates and ROAs and observe origin-validation results.
This was hands-on work, not only a presentation. That is why it deserves attention. A certificate hierarchy that looks orderly on a slide becomes a series of dependencies once someone has to publish an entity, retrieve it, validate it, serve its payload to a router and interpret the state beside an actual route. The exercise exposed defects that explanation alone might have hidden.
But the surviving evidence has limits. There are no raw configurations, packet captures, validator logs, complete version lists, failure-injection scripts or entity-by-entity results. No numerical failure rate can be reconstructed. The reports were created by entities or supporting institutions rather than by external auditors. Better completion in the second hackathon is credible evidence of repair and a more workable lab, not proof of production readiness, scale, resilience or adoption.
That distinction between having exercised a chain and having proven it under operational stress becomes central to what happened at JANOG32 on 4 July 2013.
What the room tested—and what it only feared
The JANOG32 session page and discussion record joined the reports of practical work to questions about failure. Entities discussed connecting routers to more than one cache. They raised the behaviour of RPKI-to-Router sessions, router reload, corrupt cache data and the possibility that all routes might appear Invalid. One entity proposed halting updates when anomalous results crossed a percentage-based threshold.
The page is an edited meeting record, not a verbatim transcript. It establishes that the questions were raised. It does not establish that a production network was deliberately driven into an all-Invalid state, that the proposed stop was implemented, or that JANOG adopted a rule. No numerical value for the percentage has been found in the reviewed record. There is no basis for supplying one, calling it a vendor feature, or presenting it as a Japanese production threshold.
A contemporaneous presentation, “RPKI no fukyu to kadai”, helps separate the scenarios. It mapped dependencies across loss of the RPKI-to-Router connection, router reload and convergence, corrupt cache data, and the local policy that would act on validation states. It also made a crucial observation: even a Valid route still had to pass ordinary routing policy. Origin validation checks whether the announced origin AS, prefix and maximum length are consistent with the available VRPs. It does not validate the entire AS path. It does not make a route desirable, customer-authorised, leak-free or otherwise acceptable.
The presentation is evidence that an operator-level risk analysis existed. It is not evidence that every described failure occurred, that every implementation behaved the same way, or that a controlled failure test was completed. This difference is easy to lose because a technical scenario can be described with the same vocabulary as a measurement. “What if every route becomes Invalid?” is a design question. “Every route became Invalid” is an observation. Only the first belongs to the 2013 record reviewed here.
The proposed percentage stop is similarly revealing precisely because it remained unresolved. A threshold appears attractive: if a fresh VRP set changes too much, stop distributing it before routers act on mass corruption. Yet “too much” needs a denominator and a threat model. Is the percentage calculated against all VRPs, all prefixes seen by a network, one trust anchor, one address family, one region or a previous snapshot? A globally small change could be catastrophic for one operator. A legitimate bulk update could be globally large. An attacker might deliberately trigger a breaker to freeze stale authorisations.
A repository error might sit just below the threshold.
None of those objections is a documented defect in the entity's 2013 proposal; the public evidence does not contain an evaluation. They are reasons not to transform an unanswered suggestion into policy after the fact. The important historical result is that the room recognised a control problem: when authoritative data looks anomalous, the system needs an explicit rule for whether to propagate, retain, compare, alarm or roll back. The record does not supply the final rule.
The fail-safe mechanisms were already in the protocol
JANOG32 did not invent cache redundancy or retained-data behaviour. RFC 6810, published in January 2013, already specified the RPKI-to-Router protocol before the July discussion. It allowed a router to connect to one or more caches. It described retaining data when a cache became unavailable, trying an alternate cache, and reset and update behaviour for synchronising the validated payload set.
This prior standard is more than a footnote about credit. It changes the causal account. The temporary working group can be credited with exposing implementation questions to a public operator community and relating them to hands-on experience. It cannot be credited with creating mechanisms already described in the protocol. Nor can the RFC be treated as proof that every 2013 cache and router implemented those mechanisms correctly. Standards specify expected behaviour; deployments still have versions, defaults, bugs, timers, topology and policy.
Multiple caches solve only a particular class of failure under particular conditions. A router that loses one cache can reach another if the second is reachable, sufficiently independent and serving a usable data set. Retained data can bridge a temporary interruption if its validity window and local timers make that safe. Neither control repairs a bad entity replicated through every cache. Two validators using different software may reduce dependence on one implementation, but both can consume the same inconsistent repository. Two caches in one building may share power and transport.
Two regional services may depend on the same trust anchor or cloud control plane. Redundancy is a property of failure domains, not a count on a diagram.
The protocol also cannot choose the business consequence of a validation state for every network. RFC 7115, published in January 2014 as operational guidance, made the application of validation state a matter of local policy. It urged operators to predict and measure the effect of a policy change, monitor results, and introduce treatment cautiously. Its guidance allowed NotFound routes to continue to be accepted and described staged use of validation state, rather than assuming every Invalid should immediately disappear everywhere.
Again, the date and evidence class matter. RFC 7115 is strong evidence of what best current practice said. It does not show that any named Japanese operator complied. It does clarify the governance boundary. The registry can publish certification data. A validator can decide which entities validate. An RTR cache can deliver VRPs. A router can calculate a state. The action—preference, tag, exception or rejection—remains inside the operator's routing policy. That allocation of functions leaves JANOG without authority to impose a national routing action.
This distribution of control may look like fragmentation. Operationally, it is also containment. If one policy is mistaken, it need not become everyone's mistake. If a repository has a problem, a network can distinguish NotFound from Invalid, consult its monitoring, retain or compare data where appropriate, and choose a reversible response. The cost is that safety cannot be declared at the protocol layer alone. It has to be assembled across institutions and systems.
From a registry trial to production choices
On 3 March 2015, JPNIC began a registry-linked RPKI and ROA trial environment using actual allocated resources. This was a meaningful handoff from a mock learning environment to a service connected with the registry's allocation data. It was still not automatic route-origin validation. The service material explicitly said that separate BGP router configuration was required. JPNIC could provide a certification and repository surface; an operator had to configure its network and decide what to do with the result.
Later Japanese records show convergence on many of the same failure concerns. They do not establish descent from the 2013 working group. Operators could have arrived at their designs through the RFCs, vendor guidance, RIR incidents, their own tests, global research or subsequent community meetings. A later deck's citation of JANOG30–32 shows memory and context, not causation.
IIJ's account of deploying RPKI in AS2497 is the most concrete of these outcome checks. In a JANOG47 presentation, the operator reported moving during March through December 2020 from lab and live-network tests to staged rejection across peer and upstream groups. Each router connected to two caches located at different domestic sites and using different software implementations. IIJ reported roughly 3,000 initially Invalid routes, about 0.3 percent of the full table, and a rollout spanning ten nodes and fewer than 2,000 BGP peers.
Those figures reveal the operating surface that a slogan conceals. Approximately 3,000 Invalid routes warranted investigation before rejection, but the figure did not show how many reflected harmful announcements, stale authorisations or other operational errors. The 0.3 percent denominator puts the snapshot in scale without showing that the remaining routes were automatically safe or useful. Ten nodes and fewer than 2,000 peers indicate substantial scope, but not every configuration or customer consequence. Two caches per router show deliberate redundancy, but not every shared failure domain. The deck is an operator self-report without raw configurations or an external outage audit. Its strongest contribution is the sequence: measure, investigate, divide the rollout into groups and keep infrastructure diversity visible.
The sequence also rebuts a simplistic reading of fail-safe design. Safety is not achieved only by deciding what the router should do after its cache disappears. It begins earlier, with observing how many routes would be affected and why. If an Invalid is caused by a mistaken max-length, a stale ROA, a routing error or a legitimate operational transition, simply dropping it may protect the formal authorisation while harming the intended service. Repairing data and software reduces the conflict between security and reachability.
That point is reinforced by a December 2022 APNIC case study of NTT Communications. The article reported continual monitoring of known RPKI-invalid announcements across address families and an 86.84 percent reduction in invalid announcements through software and procedures. This is a striking figure, but it remains a published case study rather than an external audit or raw dataset. It does not measure JANOG's effect. It does show how operational benefit can come from the feedback loop around validation, not merely from a final reject rule. Alerts lead to diagnosis; diagnosis leads to corrected ROAs, routing or software; better data makes stricter policy less dangerous.
Reloading nearly 800,000 routes
One fear raised in 2013 concerned router reload and the timing of validation data. A router can restore BGP state and begin processing routes while its validated payload set is still converging. If policy rejects routes marked Invalid, timing and stale state could make a restart more disruptive than ordinary BGP recovery.
A later bounded test provided counter-evidence. JPNIC's report on the JANOG50 ROV experiment described a mock environment with close to 800,000 routes. In that environment, router restart behaviour was reported as not materially different from an ordinary BGP restart. Entities also discussed redundancy, monitoring, staged handling of Invalid routes and use of local caches.
The result should narrow the fear, not erase it. “Close to 800,000” is approximate. The report is a summary, not a publication of raw timing series, configurations or every combination of router and cache. A result for one mock environment does not guarantee the same behaviour for all tables, versions, policies or failure sequences. It does show that a concern raised in operator discussion can be tested with a near-full table and may prove less severe under a defined setup than intuition suggests.
That is the useful pattern here: a concern is instrumented, tested and narrowed. The 2013 record shows the reload risk being raised; the 2022 experiment tested one version of it. Neither record warrants a universal claim. Together they show why failure-safety claims need measurements rather than folklore.
JPNIC's numbered ROV operational guideline, effective on 13 November 2024 and last updated on 27 March 2026, turns this rhythm into recommendations. It calls for monitoring cache processes, resource use and recovery of repository fetching; comparing data sets from multiple ROA caches once a day; staging rollout; and testing rollback, router restart and cache reconnection. It also describes local responses and exceptions, including use of SLURM, when the shared validation view does not safely represent an operator's circumstances, and it addresses cache disconnection that outlasts hold time.
The guideline is high-quality evidence of what JPNIC recommends. It is not a census of deployment and cannot show universal compliance. Its very breadth is nevertheless revealing. Safe origin validation is not a single configuration line. It is an operating practice involving telemetry, data comparison, capacity, reconnection, rollback, exception governance and rehearsed response. A network that enables rejection without those surrounding functions has adopted a verdict without adopting the system that makes the verdict dependable.
Two repository incidents, two missing denominators
The JPNIC disk-full event was not the first large publication fault to illuminate the chain. On 7 January 2021, a RIPE NCC repository publication inconsistency created a mismatch between parent and child certificate state. Strict relying-party implementations, particularly affected older instances with strict manifest handling, rejected all RIPE resource certificates. The RIPE NCC postmortem reported 327 relying-party instances affected and moved toward atomic publication as a remedy.
The number 327 is precise and easy to misuse. It is a count of RP instances, not necessarily 327 operators, routes, networks, customers or outages. The postmortem said the event may have resulted in outages; it did not publish a measured end-user-outage denominator. The incident took place outside Japan and proves no connection to JANOG. Its value here is as an independent observation of the same failure class: inconsistent repository state can interact with implementation behaviour so that a broad certificate set is rejected.
The JPNIC event a year later failed differently. A full disk stopped current CRL and manifest publication over 26 January to 2 February 2022. Nearly all JPNIC-issued ROAs became invalid entities. Affected routes were observed as NotFound, because usable validated authorisation data was absent. There was no published count of affected RPs, routes, traffic or users. Manual recovery restored publication. The evidence does not show that every operator used the same validator behaviour, every router received the same reduced VRP set, or any operator rejected a NotFound route.
Putting the incidents side by side prevents two bad conclusions. First, repository failure does not have a single inevitable route-state outcome. Entity validation rules, manifest handling, software versions, cache state and timing affect which VRPs survive. Second, even broad loss of validation data is not synonymous with broad loss of routing. Local policy sits between the state and the forwarding result. RFC 7115's acceptance of NotFound is one reason a loss of ROAs need not automatically remove routes.
It would be equally wrong to dismiss the incidents because user harm was not quantified. Missing impact data is an evidence gap, not evidence of zero impact. Operators still faced degraded security information, inconsistent views and possible policy consequences. The safe conclusion is specific: broad repository faults occurred; they invalidated or suppressed large data sets; relying-party behaviour mattered; and the public records did not measure the final connectivity denominator.
This is where language becomes part of engineering. If an incident report says “ROAs became invalid” and an executive summary rewrites that as “routes became Invalid”, the summary changes which control appears to have failed. If it adds “and traffic was dropped”, it invents an operator action. If it calls the result “an Internet outage”, it invents measured reachability. Precise layer names are not an escape from accountability. They are how responsibility reaches the right owner: repository operator, validator implementer, cache operator, network policy team or application owner.
The security case is the strongest objection to delay
A story focused on failure can accidentally become an argument against deploying origin validation. The evidence does not support that verdict. The strongest counterargument is that the defects visible in an early system may describe immaturity rather than a permanent safety limit—and that the security benefits grow as data and implementation practices improve.
The peer-reviewed 2019 study “RPKI Is Coming of Age” examined longitudinal ROA and BGP evidence over an eight-year horizon. Its authors found early misconfigurations had been widespread but had become very rare, and argued that the system was ready for stronger use. The dataset was global, not a measure of Japanese deployment, and its conclusions remain dependent on the authors' methods. It neither proves that RPKI is failure-free nor links improvement to JANOG. It does undermine a static inference from the 2013 hackathon defects: early roughness cannot be assumed to represent the mature system.
Origin validation addresses a real security problem. BGP ordinarily accepts an origin claim through relationships and policies that do not themselves cryptographically bind the announcing AS to a resource holder's stated authorisation. A validated ROA lets an operator detect when the prefix, origin AS or announced length conflicts with that statement. Used carefully, the signal can block or de-preference route hijacks and expose accidental announcements before they spread as trusted reachability.
The signal is limited, not trivial. A Valid result says that origin-AS, prefix and max-length conditions match a VRP. It says nothing definitive about the rest of the AS path. An attacker or leak can involve a valid origin. Ordinary prefix filters, customer-cone policy, business relationships, route-leak controls and operational judgment remain necessary. Conversely, an Invalid result is not proof of malice. It may expose stale or mistaken authorisation, a legitimate more-specific announcement that exceeds max length, or an operational change made before its ROA was updated.
That asymmetry makes monitoring valuable even before rejection. The NTT case study's reported 86.84 percent reduction suggests that software and procedure can remove invalid announcements at their causes. IIJ's staged rollout suggests that operators can investigate a reported population—roughly 3,000 at its initial snapshot—before expanding policy. The independent longitudinal study suggests that this repair work has changed the global environment over time. Together, these records support a pro-deployment conclusion with conditions: improve the data, observe the signal, stage the policy, and make reversal possible.
Failure safety is not a licence for indefinite non-deployment. Refusing to use a mature signal preserves the availability risk of fewer new dependencies, but it also preserves exposure to false origin announcements that the signal could identify. The operator's task is not to choose between perfect security and perfect connectivity. Neither exists. It is to reduce one class of risk without silently amplifying another, and to measure enough of both that a change can be defended.
Diversity, concentration and the temptation of a single service
Public cache services can lower the barrier to experimentation. They can also become concentration points. JPNIC's public RPKI cache trial began in 2015. In a notice dated 10 December 2025, JPNIC announced its retirement, citing concerns about concentration, the availability of operator and Internet exchange alternatives, later guidance, and lower use. The decision is evidence of JPNIC's own service change and rationale. It does not prove that the cache had caused outages or that every alternative was sufficiently diverse.
One visible alternative came from JPIX. In material presented at JANOG55 in January 2025, JPIX disclosed public-cache endpoints in the Tokyo and Osaka AWS regions, available over IPv4 and IPv6 on TCP port 323. Those details make the service surface reproducible: an operator can identify endpoints, regions, address families and protocol port. They do not disclose uptime, client population, underlying repository diversity, inter-region control dependencies or whether client networks use the two regions independently.
The combination tells a more nuanced story than “central is bad, distributed is good”. A central service can help operators begin, make support visible and concentrate expertise. It can also attract shared dependency. Two regional endpoints can improve geographic reach, but both may share a provider or management plane. Locally operated caches create control and observability, but they also impose maintenance burdens and can reproduce the same software or data fault across a fleet. Diversity has costs, and superficial diversity can be worse than acknowledged concentration because it creates false confidence.
A sound topology therefore names the failures it intends to separate. Physical site diversity addresses local power and facility loss. Network-path diversity addresses transport. Software diversity addresses implementation defects. Repository comparison can expose differences in validated output, although every implementation still follows the same trust hierarchy and may consume the same defective publication. Operational independence addresses simultaneous misconfiguration. No single axis contains all of them.
The public record also shows why service retirement must be analysed as continuity rather than simply withdrawal. If operators had depended on JPNIC's public cache, a safe transition required alternatives, configuration changes, tests and enough notice to avoid turning concentration reduction into a disconnection event. The notice's reference to operator and IX alternatives points toward such a transition, but no public user denominator or before-and-after audit was available in the reviewed record. One cannot infer that every former user moved safely merely because alternatives existed.
The breaker returns, still as a proposal
On 16 July 2026, a JANOG58 programme abstract returned to the old intuition in modern vocabulary. Presenters from BIGLOBE and the University of Nagasaki proposed a “VRP breaker” intended to pause distribution when an anomalous or incomplete VRP set might create false Invalid results. The mechanism was described as patent-pending. At the 20 July access date, no public deck, method, deployment result or independent validation was available.
The proposal does not complete the 2013 threshold story. It is not evidence that JANOG adopted a policy or that a Japanese production threshold exists. It does show that incomplete validated data and mass state changes remain active design questions thirteen years later. This persistence should not be mistaken for lack of progress. Systems often become mature enough to enforce only when their failure controls become more explicit.
A breaker is appealing because it introduces a moment of doubt into an automated chain. But useful doubt needs design. It must decide what baseline to compare, how to handle legitimate mass updates, whether to freeze old data or fall back to a reduced set, how long retained information remains acceptable, which operators receive an alarm, and who authorises restart. It also must avoid allowing an attacker to trigger stale-state retention. These are design questions raised by the concept, not reported findings about the 2026 proposal.
The absence of a public numerical result matters. In 2013, a entity suggested a percentage condition without an adopted value. In 2026, a programme abstract proposed anomaly-triggered distribution control without a public result at the access date. The honest continuity is the problem, not a policy lineage: operators still need a way to prevent obviously broken validation input from becoming immediate routing action. The honest discontinuity is everything else—different systems, people, evidence and maturity, with no demonstrated causal chain.
An operating model that does not fail open forever
The record from 2013 to 2026 supports a practical design position. Origin validation should become consequential, but its authority should be conditional on observable system health and reversible operator policy. What follows is analytical synthesis, not evidence of universal adoption or a proposed national rule; it makes the position concrete without inventing a universal threshold.
Reliable deployment begins with layer-specific monitoring. Repository freshness, manifests and revocation entities answer a different question from validator process health. Validator output should be measured for sudden additions, withdrawals and category changes. RTR sessions require state, serial and reconnect visibility. Routers require counts of Valid, Invalid and NotFound by peer group and address family. Routing policy requires an audit of which states change preference or eligibility. Reachability requires traffic, probe and customer evidence. A dashboard that reports only “RPKI up” hides the chain that matters.
Alarms and reports also need denominators. The IIJ deck's roughly 3,000 Invalid routes becomes more informative alongside its estimate of about 0.3 percent of the full table. The JANOG50 mock table's nearly 800,000 routes defines the scale of the bounded restart test. The RIPE NCC figure of 327 affected RP instances becomes safer when readers are told it is not a user count. JPNIC's “nearly all issued ROAs” remains incomplete without affected-route, validator, traffic and user populations. A number without its population can turn a narrow technical state into a dramatic but unsupported claim.
Diversity has to be designed around failure domains and then tested. Two caches are useful only if a router can actually switch or retain safe data when one fails. Different implementations should be compared for output and recovery, not assumed independent. Geographic separation should be tested under path and service failure. Repository inconsistency should be injected in controlled environments. Router restart and reconnection should be rehearsed with realistic table sizes and policy. The JPNIC guideline's emphasis on cache comparison, restart, reconnection, rollback and recovery gives this principle an operational form.
Policy should be staged and locally reversible. Monitoring-only deployment creates a baseline. Preference changes can reveal route-selection effects before rejection. Peer or customer groups can be introduced in controlled phases, as IIJ reported. Local exceptions such as SLURM should have owners, reasons and expiry review rather than become permanent invisible overrides. Rollback should be a tested action, not an emergency idea written after routes disappear.
Degraded trust must be separated from absent reachability. If authorisation data disappears and routes become NotFound, the immediate security posture may be weaker even while routing continues. That state deserves an alarm and repair without being called an outage. If strict validation causes route rejection, the operator needs to know whether the underlying announcement is unauthorised or the validation chain is wrong. If traffic fails, reachability evidence must be tied to the policy and path. This layered vocabulary lets operators respond proportionally.
Reversibility must not become permanent fail-open operation. Retained data ages. Exceptions accumulate. Monitoring without an enforcement plan can become ritual. A security signal that is never allowed to influence routing cannot deliver its full protective value. Operators should define entry criteria for stronger policy: stable validator output, resolved invalid populations, tested cache failure, visible rollback, accountable exceptions and a measured blast radius. They should also define exit criteria when health degrades. The criteria need not be identical across autonomous networks to be explicit and auditable.
Incident reporting should follow the chain. Repository operators can report entity and publication scope. Validator implementers can report affected versions and state transitions. Cache operators can publish service and client-impact measures where privacy permits. Networks can report route and traffic consequences. The JPNIC and RIPE NCC incidents are valuable because they disclosed mechanism and some scope; their missing denominators show what the next postmortem could improve. Better incident reporting turns one organisation's fault into shared operational knowledge without pretending that every observer suffered the same outcome.
What JANOG can—and cannot—be credited with
The temporary working group's importance lies in convening, not command. It gave entities a place to issue and inspect certificates and ROAs, build or use caches, transfer validation data toward routers, discover software defects and discuss failure. It preserved a dated public archive of concerns that remain intelligible after later tests and incidents. That is a real contribution for an operator community.
Its documented authority ended there. Resource allocation, trust-anchor operation, production certificate issuance, network regulation, JPNIC governance, router policy and national filtering all sat outside the working group's established role. The reviewed record contains neither an adopted percentage stop rule nor a documented 2013 production all-Invalid event, and it establishes no direct chain from JANOG32 to IIJ, NTT, JPIX, JPNIC's later guideline or the 2026 proposal.
Giving JANOG more credit than the record allows would also obscure the work of others. The IETF had already specified continuity behaviour in RFC 6810 and later described cautious local policy in RFC 7115. JPNIC linked certification to registry data, operated services, disclosed an incident and developed guidance. Operators built caches, monitored invalids, repaired data and staged enforcement. RIPE NCC's failure exposed publication and implementation interactions. Independent researchers measured the system's maturation. Vendors and open-source developers fixed software.
RPKI safety is a chain of contributions because no one institution owns the full path from resource authorisation to forwarding.
The absence of a single authority is not the absence of accountability. It demands more exact accountability. JPNIC can be responsible for its repository without controlling an operator's router. A validator project can be responsible for implementation behaviour without deciding business policy. An operator can be responsible for rejecting a route without issuing its ROA. JANOG can be responsible for the quality of the forum and archive without being responsible for every later deployment. Precision prevents both blame laundering and credit inflation.
Conclusion: trust the signal, rehearse its failure
The enduring image from the JANOG record is not a national switch being turned on. It is a room moving between layers: from a certificate to a ROA, from a cache to an RTR session, from a validation label to a router, and then pausing before policy. That pause was where the consequential question lived. What should an operator do if the new evidence is suddenly incomplete, corrupt or implausibly different?
By July 2013, the protocol already contained parts of an answer: more than one cache, retained data and alternate-cache behaviour. The operator discussion added concrete anxiety about reload, corrupted data and mass Invalid results. The later record added measurement. IIJ reported a diverse two-cache design and staged rollout. A near-800,000-route experiment constrained one restart fear. NTT's case study described monitoring and an 86.84 percent reduction in invalid announcements. JPNIC's guidance assembled comparison, testing, rollback and exceptions into an operating practice. Repository incidents at RIPE NCC and JPNIC showed that broad data failures were real while leaving final user impact incompletely measured.
None of that justifies an anti-RPKI conclusion. The maturation evidence points the other way. Origin validation supplies a security signal worth using, and operations improve when invalid announcements are investigated and repaired. The requirement is not to keep the signal powerless. It is to stop any single faulty layer from being mistaken for the whole truth.
A trustworthy origin-validation system therefore has to fail legibly before it can fail safely. It must say whether an entity was rejected, a VRP disappeared, an RTR session reset, a route became NotFound or Invalid, a local policy rejected it, and a user actually lost reachability. It must keep the actor and denominator attached to each statement. It must offer a tested route back from enforcement when the data chain is unhealthy, and a tested route forward when the data is sound.
JANOG's temporary experiment did not settle those obligations. It preserved a public record of them while operators were beginning to give cryptographic routing data practical weight. Thirteen years later, the proposed breaker shows the question has not vanished. The best answer is neither a magic percentage nor permanent hesitation. It is disciplined deployment: diverse, monitored, staged, reversible and increasingly willing to act as the evidence earns trust.
Sources
- JANOG RPKI Routing WG history
- JANOG32 RPKI session and discussion record
- Revised RPKI Routing WG activity report
- JPNIC Newsletter No. 55
- RFC 6810: RPKI to Router Protocol
- RFC 7115: Origin Validation Operation Based on the RPKI
- RPKI Is Coming of Age
- IIJ RPKI deployment report at JANOG47
- JPNIC report on the JANOG50 ROV experiment
- JPNIC ROV operational guideline
- JPNIC 2022 RPKI repository incident
- RIPE NCC repository inconsistency postmortem
- APNIC case study of NTT's RPKI deployment
- JPIX public-cache deployment at JANOG55
- JPNIC public-cache trial retirement notice
- JANOG58 RPKI cache-server reliability proposal
Metadata
| Field | Value |
|---|---|
| SEO title | When Trust Data Goes Dark: RPKI Failure Safety |
| SEO description | How JANOG's 2013 RPKI experiments exposed a lasting challenge: using origin validation without turning bad or missing trust data into lost connectivity. |
| Open Graph title | The Safety Question Inside RPKI |
| Open Graph description | JANOG's temporary 2013 experiment, later Japanese operations and two repository incidents show why origin validation must fail legibly and recover safely. |
| Twitter title | When Trust Data Goes Dark |
| Twitter description | RPKI makes routing safer, but repository, cache and policy failures must not be confused with one another—or with an outage. |
| Twitter card | summary_large_image |
| Focus keyword | RPKI failure safety |
| Slug | janog-rpki-failure-safety |
Featured image
| Field | Value |
|---|---|
| Alt text | Editorial diagram of validated route-origin data moving from a repository through two diverse caches to a router, with a visible pause control before routing policy. |
| Caption | RPKI resilience depends on preserving the distinctions between published entities, validated payloads, router states, local policy and observed reachability. |
| Accessibility description | A left-to-right layered illustration shows a certificate repository feeding two separately coloured validation caches. Both connect to one router, but an amber checkpoint sits between validation state and routing action. Labels distinguish entity validity, VRP availability, Valid or Invalid or NotFound route state, local operator policy, and end-user reachability. No colour alone conveys status. |
| Image provenance | Original BTW editorial illustration grounded in the publicly documented RPKI layers and failure modes cited in this article; no third-party logos or documentary photographs. |
Publication source register
- https://blog.apnic.net/2022/12/15/monitoring-awareness-and-community-at-the-centre-of-ntts-rpki-deployment/
- https://blog.nic.ad.jp/2022/7811/
- https://rpki-study.github.io/data/paper.pdf
- https://www.janog.gr.jp/meeting/janog32/doc/janog32-rpki-kimura-01.pdf
- https://www.janog.gr.jp/meeting/janog32/doc/janog32-rpki-taiji-k-okadams-yoshida-01.pdf
- https://www.janog.gr.jp/meeting/janog32/doc/janog32-rpki-yoshida-01.pdf
- https://www.janog.gr.jp/meeting/janog32/program/rpki.html
- https://www.janog.gr.jp/meeting/janog35.5/rpki
- https://www.janog.gr.jp/meeting/janog47/wp-content/uploads/2020/11/janog47_iij_rpki_20210118.pdf
- https://www.janog.gr.jp/meeting/janog52/rov/
- https://www.janog.gr.jp/meeting/janog55/wp-content/uploads/2024/11/JANOG55-RPKI-BoF_JPIX_rev2.pdf
- https://www.janog.gr.jp/meeting/janog58/pr-rpki-cache-server/
- https://www.janog.gr.jp/wg/doc/JANOG32-rpki-wg-report-11b-201403.pdf
- https://www.janog.gr.jp/wg/rpki-routing-wg/
- https://www.nic.ad.jp/doc/jpnic-01324.html
- https://www.nic.ad.jp/en/topics/2022/20220202-01.html
- https://www.nic.ad.jp/ja/newsletter/No55/NL55_all.pdf
- https://www.nic.ad.jp/ja/newsletter/No60/0240.html
- https://www.nic.ad.jp/ja/topics/2025/20251210-01.html
- https://www.rfc-editor.org/info/rfc7115/
- https://www.rfc-editor.org/rfc/rfc6810.html
- https://www.ripe.net/ripe/mail/archives/routing-wg/2021-January/004219.html

