Summary

  • RIPE NCC’s postmortem says a supplier engineer working on another fibre in the same manhole affected its fibre, disconnecting it for a couple of minutes on 27 May 2026.
  • Everything recovered automatically after the fibre was restored, but the postmortem says completion took about thirty minutes. RIPE Access, the RPKI dashboard and other login-dependent services were affected.
  • A later Q3 plan separately lists regular Keycloak backups and different traffic routes between SSO and Keycloak as work in progress. The chronology does not establish that either item was caused by the incident.
  • A small authentication-continuity receipt could record transport restoration, SSO health, representative transaction recovery and tested identity-state restoration without publishing sensitive topology or operational secrets.

Two honest clocks

RIPE NCC’s incident record is unusually useful because it keeps two durations apart. The physical interruption lasted “a couple of minutes”. After the fibre connection was restored, everything recovered automatically. Completion nevertheless took about thirty minutes. Those sentences do not contradict one another. They describe different stages of recovery.

The incident opened as intermittent availability of RIPE Access, the RPKI dashboard and other services that depend on login. An update attributed the network problem for Access to a fibre issue between RIPE NCC’s data centres. The later postmortem added the bounded physical account: a supplier engineer working on another fibre in the same manhole affected RIPE NCC’s fibre. There is no need to invent a more dramatic cause. RIPE NCC status API Incident page

The thirty-minute interval deserves the same restraint. It is not evidence that automatic recovery failed. RIPE NCC explicitly says the opposite. It is not proof of data loss, an account compromise, failed two-factor authentication or an unsuccessful RPKI action. It does not identify Keycloak as the cause. It tells us only that restoring the physical connection and completing recovery were not the same event.

That distinction matters whenever identity is a shared dependency. A fibre can pass traffic before every upstream health check, connection pool, session, cache and dependent application has reached a usable state. Conversely, a login page can answer while a privileged transaction remains unavailable. “The network is back”, “SSO is reachable” and “the requested action completed” are three claims. A postmortem that preserves the gap between them is more informative than one green timestamp.

The strongest reading is that the detail already exists

RIPE NCC may already possess the necessary operational evidence. Internal telemetry can be richer than a public status page. Route diagrams, cluster configuration, database procedures and restore results may carry security or reliability costs if disclosed carelessly. A quarterly plan is not meant to be a forensic bundle.

The organisation also deserves credit for publishing a cause at the level it could safely state, acknowledging the longer completion time and marking later work as in progress. Its service-criticality table rates RIPE Access availability as High and its confidentiality and integrity as Very High. Those ratings give a sensible reason to reveal less about the identity platform than an outside reader might like. Service criticality ratings

The case for a public receipt therefore does not begin with an accusation that engineering controls are missing. It begins with a narrower problem: the public record cannot tell a member which recovery milestone each timestamp represents. That can be fixed without exposing the machinery underneath.

A later plan is not a retroactive cause

The current Business Applications plan, last updated on 11 June, contains an SSO Improvements item. It says RIPE NCC is working on regular Keycloak backups and on different traffic routes between SSO and Keycloak to reduce timeouts. The status is “In progress”. Business Applications quarterly plan

The proximity of dates is tempting. The incident occurred on 27 May; the page was updated fifteen days later. But a nearby date is not a causal link. The route work may address a different timeout condition. Backup work concerns preservation and restoration of state, which is technically different from a severed fibre. Either project may have been planned before the incident. The published material does not say.

Archived plans reinforce the need for care. They record earlier Access work on service criticality, Keycloak and login flows, as well as monitoring and alerting work whose priority later shifted. They show an evolving programme, not a neat sequence in which each public item answers the previous incident. Archived Business Applications plans

The responsible comparison is therefore structural. The incident exposes a recovery interval. The plan names two classes of control—traffic routing and backups. A useful public record would show what each control is expected to recover and how its own test is bounded. It should not pretend that the plan explains the postmortem.

Reachability and state are different recovery objects

RIPE NCC described its 2023 Access replatforming as a move from the previous authentication backend to Keycloak running on AWS Elastic Kubernetes Service. The project was completed in July 2023, while integration and use of native functionality continued. That account explains why Keycloak belongs in the service story; it does not reveal the current topology or prove how the May incident propagated. RIPE Labs account of the Access changes

The upstream technologies clarify the categories. Keycloak’s production guidance treats multiple instances, readiness-aware routing and database reliability as separate concerns. Its database guidance makes persisted identity data crucial to availability, reliability and integrity. Its import/export documentation warns that a realm export is not a complete, automatically consistent backup: events, persisted sessions, workflow state and revoked tokens are among the omitted records. Keycloak production configuration Keycloak database guidance Keycloak import and export

Those are generic properties of Keycloak, not a description of RIPE NCC’s installation. The same boundary applies to AWS documentation. EKS provides a resilient managed control plane, but that does not prove the resilience of an application, its data, its routes or its dependencies. AWS EKS disaster recovery and resiliency

Still, the semantic split is useful. An alternate traffic route seeks to preserve or restore reachability. A backup seeks to preserve and restore state. A clustered service can answer requests while holding the wrong or incomplete recovered state. A perfectly restored database can remain unreachable. The controls may complement one another, but they cannot be compressed into a single word such as “redundancy”.

Access is a gateway to consequential actions

RIPE Access is not merely a convenience login for a brochure site. RIPE NCC describes it as the single sign-on gateway used by member services. Its published authentication and security-key policy covers the LIR Portal and the RPKI dashboard. RIPE NCC Access policy, ripe-843

The current RPKI Certification Practice Statement says the Online CA relies on the SSO mechanism used by the LIR Portal to identify authorised requesters. The RIPE Database documentation likewise says SSO credentials managed by RIPE NCC Access can authorise web updates to protected objects. RPKI CPS, ripe-851 RIPE Database authorisation model

That does not mean the May interruption changed a route, a ROA or a database object. No such outcome is in the incident record. It means that the recovery claim should be made at the correct layer. An identity gateway can be healthy as a service while a dependent transaction is still failing. A dashboard can load while a privileged action is not yet confirmed. The public receipt needs one or two safe representative transactions, not a universal claim about every member operation.

Authority must also stay separated. RIPE NCC operates Access and its member-facing services. A cloud provider operates parts of the underlying platform. A supplier may work on physical infrastructure. Members control their own accounts, local applications and decision timing. An upstream documentation page can define technology behaviour; it cannot attest to a particular RIPE NCC recovery. A receipt is valuable precisely because it joins these boundaries without erasing them.

What an authentication-continuity receipt would contain

The smallest useful record begins with the incident identity and four clocks: physical impairment detected, transport restored, SSO declared healthy and representative dependent transaction confirmed. If services recovered at different times, the receipt keeps separate rows instead of choosing the latest or most flattering one.

For the route control, it records the intended failure boundary, the health signal, the failover trigger, whether failover was automatic or manual, the observed switch time and the safe test transaction. It does not publish path diagrams, addresses, credentials or thresholds that would make abuse easier. “Independent path” should be used only when independence has actually been tested against a named failure domain.

For state restoration, it records the protected data classes, backup frequency, encryption and retention policy at a publishable level, the restore-test date, the tested software version, the recovered boundary, measured recovery point and recovery time, and any deliberately excluded state. A statement that a backup completed is not a restore result. A realm export is not automatically a database recovery. The receipt can say “details restricted” while still naming the test and its outcome.

The record then adds ownership and lineage: which team made each assertion, which evidence is internal, which independent review was used, what remains unknown, and how a later correction or supersession links back. It expires. A successful test from one version and one topology does not become a permanent property of the service.

This is not a demand for a new certification authority. It is a bounded operational receipt. RIPE NCC can publish it alongside the status incident or quarterly-plan item, keep sensitive evidence under controlled access and correct it without rewriting history.

Recovery is complete when the claim says what completed

The most valuable sentence in the postmortem may be the plainest one: recovery was automatic, but completion took about thirty minutes. It resists the usual compression of a layered system into a single uptime number.

The next improvement is equally modest. Give each layer its own object and clock. State what the alternative route protects. State what a restore test preserves. Name the representative transaction that closes the incident for a member. Preserve the uncertainty where the public record cannot go further.

RIPE NCC need not publish its architecture to make continuity legible. It needs only to stop one recovered layer from silently speaking for all the others.

Sources

  1. RIPE NCC status API: 27 May 2026 incident
  2. RIPE NCC status page: intermittent availability of RIPE Access
  3. RIPE NCC Business Applications quarterly plan
  4. RIPE NCC archived Business Applications plans
  5. RIPE NCC service criticality ratings
  6. RIPE Labs: Access replatforming and security changes
  7. RIPE NCC Access policy, ripe-843
  8. RIPE NCC RPKI Certification Practice Statement, ripe-851
  9. RIPE Database authorisation model
  10. Keycloak production configuration
  11. Keycloak database guidance
  12. Keycloak import and export
  13. AWS EKS disaster recovery and resiliency