Summary

  • RIPE NCC marked its 22 August Alfresco maintenance complete at 17:56 CEST, more than two hours before the scheduled window ended.
  • A 24 August incident said document creation remained operational while some documents could not be viewed or downloaded; it was resolved after about one hour.
  • A second incident opened on 26 August, reported processing delays, moved to monitoring after a configuration update on 27 August and was then resolved.
  • The records establish sequence and a shared system, not causation. What is missing is a public lineage joining change acceptance, operation-level tests, recurrence and case reconciliation.

Three green states tell three different stories

The maintenance record is admirably precise about the planned interruption. RIPE NCC said it would upgrade Alfresco in production from 08:00 to 20:00 CEST on Saturday, 22 August. Alfresco stores member-related documentation. During the window, the LIR Portal and document system would be unavailable, and membership applications, resource transfers, resource requests, merger-and-acquisition requests and other ticket-created actions could not be completed.

At 17:56, the status page posted a short conclusion: the scheduled maintenance had been completed. The affected LIR Portal component moved from maintenance to operational. That is one legitimate kind of closure. It says the scheduled operation ended and the named component was returned to service.

Forty-five hours later, a different record opened. RIPE NCC said document creation remained operational, but some documents might not be available for viewing or download. Resource transfers, resource requests, merger-and-acquisition work and any action for which a ticket was fetched were affected. The incident moved from partial outage to operational at 16:00 on 24 August, about an hour after it began.

A second record with the same incident name opened on 26 August. This time the public description did not repeat the create-versus-retrieve distinction. It reported delays processing documents related to all request types, merger-and-acquisition requests and transfers. The following morning, RIPE NCC said it had updated the Document Management System configuration and was monitoring the result. At 16:21 on 27 August, it marked the incident resolved.

All three records eventually turn green. They do not mean the same thing. Completed closes a planned change. Resolved closes an observed incident. Operational is a component state. None, by itself, identifies which document operations were tested, how many cases were affected, whether queued work was replayed or whether a later incident is considered a recurrence of an earlier one.

“After” is not “because of”

The sequence invites a tempting story: an upgrade was completed, retrieval failed, and a configuration change fixed it. The official record does not establish that chain of causation.

The status entries do not publish a root cause for the 24 August failure. The later entry says configuration was updated, but it does not identify the changed setting, connect it to the Saturday release or say whether it corrected the same mechanism as the first incident. The two incidents may share a cause, may expose different faults or may be unrelated beyond occurring in the same system. A responsible article cannot choose among those possibilities.

That boundary does not make the chronology trivial. In operational governance, uncertainty is itself a state that should be recorded. A change record can say causal link: unconfirmed. An incident can say recurrence: under review. A resolution can distinguish service restoration from root-cause closure. Such fields prevent both extremes: quietly treating connected events as unrelated, or turning proximity into accusation.

The first incident also demonstrates why a single health label is too coarse. Creating a document and retrieving it are separate operations. A write path can succeed while a read, view or download path fails. A portal can answer while a case worker cannot fetch the evidence needed to complete a request. A component-level green state cannot describe those differences unless the acceptance record names the operations it exercised.

Completion needs an evidence boundary

The point of an acceptance test is not to promise that software will never fail. It is to say what was checked before the institution declared the change complete.

For a document system used in registry administration, that boundary can remain compact. It can record the released build and configuration fingerprints; a test that creates a document; a test that retrieves, views and downloads it; permission checks for representative roles; a ticket fetch; and confirmation that an existing case can proceed without corrupting or losing its evidence. The public version need not expose filenames, member identities or internal topology. It can publish operation classes, sample counts, pass times and the identifier of the fuller protected record.

The observation window matters too. A Saturday maintenance may pass a small immediate test and still encounter a different load or workflow when staff and members return. That does not make weekend scheduling wrong. It means maintenance complete and post-change observation complete should be separate states. If an incident arrives during the observation window, the record should link them even when causation remains undecided.

RIPE NCC already supplied the raw ingredients for such a model. Its first incident separated create from view and download. Its second exposed a configuration intervention and a monitoring interval. Its status system preserves maintenance and incident identifiers. The missing item is the join.

Reconciliation is the part readers cannot see

Restoring retrieval is not necessarily the end of a member workflow. A transfer may have waited for a document. A resource request may have remained in a queue. A merger case may have been opened but not advanced. An automated retry may have succeeded; a manual action may still be needed.

The current records name affected classes but give no bounded account of the work left behind. They do not need to reveal whose transfer was delayed or what a document contained. A privacy-safe closure could state how many cases were in scope, how many required replay or manual review, how many were reconciled and whether any remained unresolved at the time of closure. A zero is useful evidence. So is not yet known if it later receives a final value.

This is where the doctrine of the registry as a narrow ledger becomes practical rather than rhetorical. RIPE NCC does not need a grand claim of infallibility. It needs a reconstructable account of the administrative evidence on which resource decisions depend. When the document surface fails, the important question is not whether the institution’s reputation stayed green. It is whether the records and pending decisions were preserved, recovered and checked.

One continuity record can join the sequence

A useful public control would bind the three records without pretending to know more than they say. It would carry:

  1. the maintenance identifier, release and configuration fingerprints;
  2. the create, fetch, view, download and case-processing tests run at completion;
  3. test counts, completion time and the planned observation window;
  4. the two later incident identifiers and affected operation classes;
  5. causal-link and recurrence states, including unconfirmed where appropriate;
  6. configuration or rollback references without sensitive implementation detail;
  7. backlog, retry and case-reconciliation totals; and
  8. the time at which service restoration, reconciliation and root-cause review each closed.

That record would not convert every defect into a governance scandal. It would do the opposite. It would keep technical facts narrow, separate and correctable. Readers could see that a scheduled change ended, that two later failures occurred, that the causal relationship remained open or was later resolved, and that affected work was accounted for.

The present public record supports a limited conclusion. RIPE NCC completed a planned document-system upgrade, then recorded two retrieval failures in the following working week. Both incidents were resolved. No affected member, document or resource has been identified, and no causal link has been published. The remaining accountability gap is not proof of a hidden failure. It is the absence of a receipt showing what “complete” tested and how later recurrence was reconciled with it.

Sources