Summary

  • On 23 December 2019, a storage failure combined with a configuration error defeated ARIN’s virtual-machine failover. Its public website and ARIN Online returned at different times through a separate disaster-recovery path.
  • ARIN’s Board and management then ordered technical reporting, remediation plans, and better service visibility. Those records show a more observable response, but they do not independently prove that every risk was eliminated or a service-level target was met.

The hour when redundancy stopped being a promise

At 12:35 p.m. Eastern time on 23 December 2019, ARIN Operations received several monitoring alerts about virtual machines that supported customer-facing services. Staff tried to force the machines onto standby hardware. The failover attempt did not work. By 12:55, the virtual machines had lost their hosting in that part of the network; ARIN’s website and its ARIN Online customer application were unavailable. Those are the two surfaces the public incident account names.

It does not provide a service-by-service statement about every registry function, so the account should not be enlarged into a claim that all ARIN services—or the underlying registration records—were down.[4]

The details matter because availability is often discussed as though it were a property of a whole organization. Customers, however, encounter separate services: a public website, an account application, registration workflows, directory lookups, routing-security tools, notifications, and support. One may fail while another remains usable; two may share a dependency that neither user interface reveals. A statement such as “the registry was available” is too broad to tell a network operator what action was possible. So is the opposite statement if it implies that every registry system failed.

ARIN’s published timeline gives a rare view of that distinction. It identifies when monitoring first raised concern, when the initial standby path failed, when staff chose another recovery option, and when the two named services returned. It does not present a complete dependency graph, an availability series, or a list of systems that were unaffected. The appropriate reading is therefore narrower and more useful: the incident took down the website and ARIN Online, and the record leaves the status of other services unestablished.[4]

For John Curran, then ARIN’s President and CEO, the story is not that he personally fixed a storage appliance. The operating account was written by Chief Operating Officer Richard Jimmerson and attributes alert handling, troubleshooting, vendor coordination, restoration, and the after-action review to Operations and staff. Curran’s documented role appears at the management and Board boundary: he supplemented the report to trustees, participated in the follow-up record, and later had to make system condition legible as an executive accountable for operations.

That boundary lets the article examine leadership without turning an organizational recovery into a solo-hero narrative.[1][2][4][5]

A different outage had already raised a different question

The December failure was not the only service interruption in ARIN’s recent record. At a January 2019 Board meeting, Curran described a recent outage of the ARIN.NET domain for parties performing DNSSEC validation. He said he had reviewed the post-mortem with trustees and would meet with Chief Technology Officer Mark Kosters to align on ARIN’s mission-critical systems. The Board chair requested a report after that work.[3]

The minutes do not give the DNSSEC incident’s duration, root cause, number of affected parties, or a completed report. They cannot support a comparison of its severity with the later storage event. They do establish a distinct operational concern: a domain-name validation problem can prevent a validating party from reaching a domain even when the immediate issue is not the same as an application outage. That is a different failure mode from the December cluster loss, and the article keeps the two apart.

The pairing is useful because it shows why “the internet” and even “ARIN’s services” are not single uptime objects. In one record, the affected surface is the ARIN.NET domain for DNSSEC validators; in the other, the published account names the website and ARIN Online application. The documents do not establish whether those incidents shared infrastructure, people, or causes. They should not be joined into a single incident narrative simply because the same CEO briefed the Board.

The January discussion also matters for Curran’s role. The minutes show him taking the post-mortem to the Board and planning a meeting with the CTO about critical systems. They do not say he wrote the post-mortem, performed the technical diagnosis, or personally implemented a control. A leader’s accountability is visible in whether a technical finding enters a management process, receives an owner and follow-up, and returns in a form the governing body can evaluate. The meeting record proves some of that handoff was requested; by itself it does not prove how completely it was closed.[3]

What failed in December—and what did not

The public account by Jimmerson says the affected virtual machines ran on a cluster designed for high availability, with redundancy throughout the system. When staff tried a manual failover to other hardware on standby, that path failed as well. Troubleshooting with the virtualization vendor and rebuilding nodes pointed to access to the shared-storage platform as the likely culprit. The storage vendor could not resolve the problem quickly enough, and the shared system remained unavailable to the virtual-machine cluster.[4]

After an after-action review, ARIN described the cause more specifically: a hardware failure coupled with a system misconfiguration on the storage appliance caused the virtualization engine to fail. The hardware was replaced, vendor support validated the corrected configuration, and ARIN reported normal operation. “Hardware failure” alone would hide the configuration element. “Misconfiguration” alone would hide the failed component. “Vendor issue” would transfer too much responsibility away from the operator that designed, configured, and depended on the system. The published explanation treats the event as a coupled failure.[4]

That is the central technical lesson, and it is more precise than the familiar warning that redundancy can fail. Redundancy protects against the failures it actually separates. If primary and standby machines depend on a storage service that is not independent, a problem in that shared service can disable both. The attempted failover can be correctly initiated and still fail to create a working application. A standby host is not a disaster-recovery plan merely because it is powered on; the data path, configuration, authentication, network path, operational procedure, and staff access must also work under the relevant fault.

ARIN’s account says the virtual machines lost hosting in one part of its network, then describes a separate disaster-recovery site to which the website could be swung. That separate path eventually restored the site. It would be inaccurate to say that the primary high-availability system worked as designed: the attempted failover did not. It would also be inaccurate to say that ARIN had no recovery capacity: the disaster-recovery site was maintained and tested, and it provided an alternate route to restore the website. Both statements can be true at once.

Resilience is not a binary label; it is a chain of specific paths that work—or do not—under particular failures.[4]

The record is candid about one important unknown. It names the website and the ARIN Online application as unavailable, then describes restoring those surfaces. It does not tell readers whether Whois, RDAP, authoritative or reverse DNS, the Internet Routing Registry, RPKI repositories, provisioning interfaces, or the data behind customer records were affected. The article therefore does not infer a broad data-plane failure from a web application incident, and does not infer that every other function stayed healthy from their absence in the account.

A service map and an incident-specific status record would be needed to answer those questions.[4][9]

Recovery occurred on two clocks

At 3:30 p.m., after the outage had lasted and the repair time remained uncertain, staff decided to move ARIN’s website to its disaster-recovery site. The website was restored by 4:00 p.m. ARIN Online returned at 5:10 p.m. The report gives separate restoration times rather than one generalized “back online” timestamp. The difference—70 minutes between the site and customer application—shows that even related public services may have distinct dependencies and cutover steps.[4]

That separation has practical meaning. The website can deliver public explanations and instructions. ARIN Online is the customer application. Restoring the first does not establish that a customer can use the second, and restoring the second does not establish every registry endpoint is healthy. If a network operator needs to submit a request, manage an organization, or check an account, the relevant recovery clock is the application’s, not the homepage’s. If the operator needs a routing-security publication or a registration lookup, another clock may matter.

The 2019 report does not give those additional clocks, so the later public service taxonomy is helpful context but not retrospective evidence about that day.

The decision to use the disaster-recovery site also shows a threshold judgment under uncertainty. Staff had attempted troubleshooting and vendor coordination; by 3:30, the duration of the incident and unknown time to repair made waiting for the original environment less attractive. The record does not say how the threshold was calculated, who had authority for each step, or whether the decision met a prewritten recovery-time objective. It does show that ARIN had a tested alternate site and that the team chose to invoke it when repair time could not be bounded.[4]

The rest of the restoration was deliberately staged. The storage vendor continued work through the Christmas holiday. The failed component was replaced by 2 January, after which the affected cluster returned to operation. Moving the website back to the primary data centre would itself require an outage, so ARIN paired that change with a scheduled maintenance window on 25 January rather than adding an unplanned cutover. This is not the same as saying the system was fully remediated by 2 January: that date marks the component replacement and return of the impacted cluster, while the separate site transition was still scheduled.[4]

These distinctions are often lost in post-incident summaries. “Service restored” can mean that customers have a working alternative, that the faulty component is replaced, that the primary environment is back in use, that the defect has been corrected, or that all corrective actions are complete. Those are different states. ARIN’s timeline gives enough detail to distinguish several of them; its January-to-May Board record adds more. A good reliability account should preserve that granularity instead of compressing every milestone into one reassuring sentence.

The CEO’s authority was broad, but the incident work was distributed

ARIN’s current organization pages describe the President and CEO’s place in the institution. The Board maintains authority over ARIN’s scope and mission, and the Board and CEO establish strategic direction and fiscal oversight. The CEO and staff execute strategy through operational management. The Board selects the President; the President appoints and supervises operations staff, serves as a voting Board member, and liaises between the Board and Advisory Council.[1][2]

Curran’s biography adds that he had been a CTO at BBN/GTE-Internetworking, XO Communications, and ServerVault; led development of early research networks; and provided technical leadership in BBN’s transition to commercial Internet services. He was a founding ARIN trustee in 1997, chaired the Board through early 2009, and became President and CEO in 2009.[1] That experience makes a profile of him relevant to infrastructure operations. It does not erase the separation between a CEO’s institutional responsibility and the work of staff diagnosing a storage failure.

The Board pages also describe the President as the tenth, ex-officio trustee in a ten-member Board, alongside nine trustees elected by ARIN general members. That arrangement creates a role with two kinds of responsibility: executive management of the organization and participation in the body that sets its scope and oversight. In January 2020, the minutes identify the Board as the body that requested remediation and infrastructure reports; they identify the COO as the person who presented the outage report and technical updates.

Curran supplied additional context, served as a conduit for some documents, and later discussed the way reports should be made more useful.[1][5][6][7]

This is a better way to evaluate the leadership record than asking whether a CEO “caused” an outage. A shared-storage fault and configuration error are system conditions. The institution’s response includes detection, authority to switch services, vendor management, customer communication, corrective engineering, budget and staffing decisions, risk review, and evidence to the Board. Some steps belong to Operations; others belong to executives or trustees. A careful article can follow those interfaces without assigning a technical act to a person the source does not name.

The public sources are also uneven by design. The January 2019 minutes are an abbreviated Board record. Jimmerson’s March 2020 blog is a first-person operational account with a timeline and after-action explanation. Later minutes record management reporting and Board requests, not every engineering artifact. The current staff page defines responsibility but is not an incident log. Reading these genres separately prevents two opposite mistakes: treating a CEO’s title as proof he personally executed an incident response, or treating an operational post-mortem as though it answered every governance question.[1][2][3][4][5]

The Board turned an incident into a reporting sequence

At its meeting on 22–23 January 2020, the ARIN Board received the December outage report. The minutes say the COO provided the reason for the incident, while the President supplemented the information. The chair thanked both for candor and for explaining the failures that preceded the outage. The Board then asked for four follow-ups: a technical-infrastructure report by April; an accelerated account of what was needed to move long-term remediation forward; priority for strengthening the disaster-recovery profile; and an updated incident assessment at the April meeting, including information that would let trustees understand downtime risks.

The minutes say the Board would assess the matter after receiving the material.[5]

The President also said a report would be published for the community. The operational account appeared on 19 March, authored by Jimmerson. This chronology matters: public disclosure and internal oversight were related but not identical. A customer-facing post explains what ARIN chose to disclose about the event. A Board report can contain an infrastructure inventory, risk material, a remediation timeline, and other internal management details. One should not be mistaken for the other.[4][5]

The requests were concrete enough to create checkpoints. The technical-infrastructure report had an April target. The Board asked to move long-term measures faster and explicitly elevated disaster recovery. The incident assessment was supposed to include downtime risk rather than only a narrative of what happened.

That shifts the question from “Did the website come back?” to “What dependencies were involved, how much outage could recur, and which controls change the likelihood or impact?” The record is stronger when it captures those questions, because a post-mortem without an owner, timeline, or risk model can be informative yet operationally inert.

By 25 March, the Board had a short progress update. The COO said a fuller update would be provided in April, documents were forthcoming from the President, and the remediation deadline was the end of 2020. Staff was preparing a document for the Board showing how the work would be done; the COO said the plan was on schedule.[6] This is evidence that management had a path and date in March. It is not evidence, on its own, that every task later met the date.

On 21 May, the minutes record a more substantial response. The COO had completed the outage-remediation report and a technical-infrastructure supplement to the technical-debt review. The infrastructure summary responded to the Board’s January request. An outage-remediation requirements document described the proposed path to finish before the close of 2020. Trustees wanted reports to connect systems to business functions and to include start and end dates so that progress and possible acceleration could be seen over time. The Board asked that the technical reports continue quarterly.[7]

Those documents mark a shift from a single incident explanation to a reporting system. That is an observable management result, even if the public minutes do not expose the entire technical package. A recurring report can make aging infrastructure, dependencies, ownership, and dates visible to trustees who otherwise see only a summary. It can also fail: a dashboard can compress a risk into a green cell, a deadline can slip, and a statement that work is “on schedule” can be true while the remaining exposure is still material. The evidence supports the existence and shape of the reporting loop; it does not certify every control within it.

Technical debt made a reliability question legible to the Board

The May minutes place outage remediation beside technical debt, a useful pairing because equipment refresh is not merely a software housekeeping concern when a shared component supports customer-facing applications. The Board discussed how technical reports should describe which business functions depended on which systems, how a proposed change affected ongoing operations, what costs might be saved, and when work would begin and end.

Curran said he was developing principles and timelines for considering outsourced technical solutions; the chair welcomed a written framework to help trustees understand why outsourcing could or could not be used.[7]

That exchange makes the leader’s contribution more specific than a general commitment to reliability. Curran was not reported as choosing a replacement storage array or leading a failover. The minutes do record him shaping how system condition should reach the Board: link technology to business processes, supply a time-bound trajectory, and state decision criteria around outsourcing. Such a reporting choice can change what oversight is possible. If a Board sees only product names and age, it cannot easily compare cost, business exposure, and transition risk.

If it sees which service depends on a system, what failure can interrupt, and what mitigation is funded, it can ask more consequential questions.

There is an important limit. The public version of the May minutes says the completed outage report and technical-infrastructure summary went to the Board, but the minutes do not reproduce their full contents. The public record is therefore strong on process and weak on the underlying engineering detail. A reader can verify that the requested documents were reported completed and that quarterly follow-up was expected; a reader cannot infer the exact storage architecture, failover-test results, or completion of every remediation task from that sentence.

That distinction should survive into the article rather than being replaced with an optimistic or cynical guess.[7]

The Board’s later reporting agenda widened from one incident to reliability as a management property. At its February 2021 meeting, trustees sought less historical reporting and more useful measures of system performance and customer service. Curran was asked to look at metrics, including possible service-level reporting for major services. The minutes describe discussion of a framework for availability, reliability, consistency, and privacy, and a requested map of dependencies that would identify crucial systems.[8]

This language exposes what a service objective must contain. “Availability” is not a score unless the service boundary and measurement window are defined. “Reliability” needs a failure mode and a recovery expectation. “Consistency” may concern whether separate information surfaces agree. “Privacy” can constrain where operational visibility is exposed. A service-dependency map links those abstractions to the applications and infrastructure people use. The record shows trustees asking for this kind of management view; it does not supply a public, numeric recovery-time objective for the 2019 incident.

The same February 2021 minutes say ARIN was up to date on many systems at the close of 2020, with the original goal to bring many systems out of a “red” state. The Board asked how to avoid technical debt returning and requested metrics that would be meaningful to customers and trustees. A changing set of technical-debt charts is not itself proof of resilience, but the question points toward an important operational discipline: maintenance must have a recurring budget and owner, not just a crisis-generated project. Whether that discipline was fully implemented is beyond what these minutes establish.[8]

Visibility is a control, not a substitute for recovery

On 5 April 2021, ARIN announced a public Service Status page. It listed several service families: ARIN Online, Reg-RWS and provisioning; Whois, RDAP, RPKI, IRR and other registry services; mailing lists; reporting services; and the website. Readers could subscribe by email, text, Slack, or other methods, and the announcement said the page fulfilled ACSP 2020.5, a Service Availability Page proposal.[9]

The status page solves a different problem from a disaster-recovery site. Recovery provides a path to restore service. Status communication tells users which service is affected, what the operator knows, and where to watch for updates. The status taxonomy helps readers avoid treating a public homepage as a proxy for every registry function. Subscription mechanisms help a network operator learn about an incident without continuously polling the website that may itself be affected.

Neither function proves the other. A beautifully structured status page can report an outage without restoring anything. A tested recovery site can restore a website while customers do not know that a separate application remains unavailable. A full-service map can make dependencies clearer but does not make them independent. A public status dashboard is also not an SLA: it does not necessarily publish a target, measurement method, error budget, recovery objective, or credit. ARIN’s announcement describes visibility and notification, not those contractual metrics.[9]

It would also overstate the record to say that the status page was created solely because of the December incident. The announcement says it fulfilled a member-community proposal, ACSP 2020.5, and is authored by CTO Mark Kosters. The public sources establish that the page appeared after the outage and that it made service categories and updates more accessible. They do not establish a direct causal chain from the storage failure to that specific product decision. Sequence is not causation, and chronology should not be drafted into a hero story about the CEO.[9]

Still, the page is a notable change in how accountability can be observed. Before it, ARIN said service-impacting notifications were often communicated through the ARIN-announce mailing list and meeting updates. The status page consolidated reporting and allowed subscriptions across channels. It created a place where outages, maintenance, and other service changes could be seen separately from policy discussions. That makes a consumer-facing operational claim more testable—provided incident entries preserve service names, impact, timestamps, and resolution rather than becoming a stream of vague green checks.[9]

For network operators, the usefulness is immediate. During an incident, an engineer may need to know whether a ticketing application, routing registry, certificate publication service, or information page is unavailable. If each is reported as a distinct component, the operator can select a fallback: queue a request, rely on cached data, delay a change, or check an independent source. If the page only says “ARIN service” is operational, it does not answer the question. The taxonomy ARIN announced gives a better starting vocabulary, though the article does not claim every historical event is mapped retroactively to every category.

What the later record does—and does not—close

In August 2022, the Board discussed technical-debt and risk reporting. When a trustee referred to a previous outage involving NetApp, the COO said that changes to leadership structure and infrastructure upgrades had mitigated the prior high-risk area, and that ARIN had moved to solutions from other vendors for better tools and more budget-friendly pricing. The COO added that this change was tracked in a management report rather than the technical-debt report under discussion. The Board also discussed making charts more intelligible around ARIN’s main applications, and a new risk register with mitigation strategies.[10]

This is meaningful evidence of later change: the operations executive reported infrastructure upgrades, vendor changes, and mitigation. It is not an independent audit. The minutes do not say every storage risk had vanished, nor do they enumerate the status of each 2019 follow-up. They do not say the later vendor decision was made only because of the 2019 incident. Treating the COO’s report as proof of every engineering outcome would give a Board summary more evidentiary weight than it contains; dismissing it would be equally unsound. It is management testimony recorded in Board minutes, and should be identified as such.[10]

ARIN’s 2025 annual report lists as strategic objectives providing industry-leading registry services, strengthening organizational practices for consistent service, advancing routing security, and increasing outreach to Caribbean members. It identifies Curran as President and CEO and the tenth Board trustee, alongside nine members elected by general members.[1][11] These current statements show that service consistency remains part of ARIN’s public operating agenda. They do not report the uptime of each service, the test results of the 2019 disaster-recovery plan, or a long-term success metric attributable to Curran.

The public record therefore supports a measured conclusion. ARIN disclosed a specific outage with its own sequence, causes, and recovery times. The Board converted it into requests for infrastructure information, remediation, and downtime-risk assessment. Management reported progress over the following months; the technical-debt discussion became more regular and tied to business functions. A service-status page later gave users a clearer map of service families and a notification path. By 2022, the COO reported that infrastructure upgrades and vendor changes had mitigated the earlier high-risk NetApp area.

Those are real steps in an accountability chain.[4][5][6][7][8][9][10]

But the same record does not show a public before-and-after service-availability chart, an independently verified control test, a complete map of the 2019 dependency chain, or a public closure checklist tied to every January task. That absence in the sources reviewed is not proof that ARIN lacked such records or failed its remediation. It means a reader cannot responsibly claim more than the published evidence demonstrates. A reliable profile should say where the trail grows stronger and where it stops.

The operational lesson is service-specific

The 2019 episode is often compressed into a familiar narrative: redundant systems failed, then a disaster-recovery site saved the day. The compressed version misses the useful questions. Which service did the first failure interrupt? Which recovery attempt depended on the same storage? What did the separate site restore? How did the application’s return time compare with the website’s? What was the after-action cause? What remedy was assigned, by whom, with what date? Which record later showed whether risk changed?

ARIN’s public materials answer some of those questions. They name two unavailable surfaces, a failed standby attempt, a tested DR site, restoration times, a combined hardware/configuration cause, and dates for replacing the component and returning the cluster. Board minutes show required follow-up and a management reporting cadence. The service status announcement shows distinct public service categories. Other questions remain open in the documents: the status of unlisted registry services on the day, actual recovery-time targets, test evidence for each failover path, and a public closure record for every remediation action.[4][5][7][9]

That method is useful beyond ARIN. A registry is relied on by network operators but is not a single application. Its staff administer policy and records; its services present those records and tools to customers; its authoritative data may feed other technical systems; and its governance body allocates resources and oversight. A disruption in one surface should not be treated as proof of a disruption in every layer. Nor should the persistence of a record be confused with the ability to reach a service, submit a request, or learn which state is current.

For customers, the distinction informs contingency planning. Operators can identify which workflows need current registration data, which ones can use cached information for a time, and which changes should pause if a provisioning service is down. They can subscribe to incident notifications and record the status source used for each decision. They can also ask providers for concrete service boundaries rather than a single uptime promise. None of this assumes that the operator’s own business can tolerate a registry outage; it asks which activity is affected and what alternative remains.

For ARIN, the lasting question is whether reporting links incidents to measurable service objectives. A quarterly technical-debt report can surface an aging platform. A dependency map can show what depends on it. A disaster-recovery exercise can test whether an independent path works under a realistic failure. A status page can tell customers which service is affected. A post-incident report can compare actual restoration with an objective and list follow-up actions. Each artifact answers a different question. If one is missing, the others should not be used as a substitute.

That is where a profile of Curran’s leadership is most grounded. His documented work through the event was executive and institutional: report to the Board, pass on information, propose a path for remediation and system oversight, and help make the operational picture more usable to trustees. The technical response belonged to Operations and staff. The CEO’s result is not measured by whether he personally touched a console; it is measured by whether the system gave the Board and service users enough reliable information to understand failures and make decisions.

The sources show improved reporting and visibility, but they stop short of proving a complete reliability target was achieved.[5][6][7][8][9]

A stronger claim requires a stronger receipt

The central distinction is between four statements that are often collapsed: a service has returned; a component has been replaced; a known defect has been corrected; and the risk of recurrence has been reduced to an agreed level. ARIN’s report documents the first three for parts of the December sequence, though “normal operation” is the organization’s own description. The Board minutes document a remediation process and risk reporting. The later COO statement reports mitigation of a high-risk storage area. The reviewed sources do not provide an independent quantitative measure of the fourth statement.[4][5][6][7][10]

The distinction is not pedantry. If a customer’s only evidence is a green status lamp after recovery, the customer cannot tell whether a fragile configuration remains. If a trustee sees only a narrative post-mortem, the trustee cannot compare risk before and after a funded change. If management reports only that an old platform is “up to date,” it may not reveal whether the service path was exercised against the failure that caused the outage. Good evidence identifies the service, records the dependency and fault, shows restoration time, specifies the corrective action, and returns with a test or measurement that demonstrates what changed.

ARIN’s records move in that direction over time. The December 2019 post describes chronology and cause. The Board asks for infrastructure and outage-risk detail. Quarterly reporting adds technical debt and business-function relationships. A 2021 discussion asks for service-performance measures, availability and reliability framing, and a dependencies map. The public status page separates service families. In 2022 the COO reports mitigation and vendor changes. This is an accumulating public trail, not one definitive proof.

It invites the next question: which measure would let a member verify that the recovery path for each named service can survive its own shared dependencies?[5][6][7][8][9][10]

Curran’s long tenure matters here because operational accountability accumulates. He joined ARIN’s founding Board in 1997, chaired it until 2009, then became its President and CEO. The current organizational structure gives the CEO operational authority, while the Board retains scope, mission, and oversight responsibilities. That is not a personal mandate to guarantee that no system ever fails. It is a durable position from which an executive can require the organization to define service boundaries, fund recovery, disclose incidents, and return evidence to the Board.[1][2]

His earlier technical leadership at BBN and other Internet companies provides context, but the public record should not be used to infer a private motive or claim he personally designed ARIN’s storage cluster. The more observable line is enough: after two different service interruptions, the Board asked for formal review; reports were produced; technical debt and dependencies received recurring attention; and service status became more legible to customers. That narrative is modest, but it is stronger than the claim that an experienced CEO “saved” the system. It follows the evidence instead of reputation.

Conclusion: resilience is something a service can show

On 23 December 2019, a cluster built with high availability lost the service of its shared storage, and the standby path did not restore it. ARIN’s operations team switched the website to a separate recovery site and restored the customer application later. The event was technically specific, operationally bounded in the report, and followed by Board-directed remediation and reporting.[4][5]

John Curran’s part in the record is similarly bounded. He was the President and CEO with an executive role on the Board; he supplemented the December report, delivered documents into the Board process, and was later tasked with improving performance metrics and explaining dependencies. The COO and operations staff described the diagnosis and recovery. Later management reports recorded technical upgrades, a shift in vendors, and a public service-status facility. None of those facts should be inflated into a claim of personal troubleshooting or a guarantee of future uptime.[1][2][5][7][8][9][10]

The more durable lesson is that a registry’s public authority and its operating reliability are different questions. A website, customer application, DNSSEC validation path, registration database, and RPKI publication service are not interchangeable. The user needs to know which one failed, what its recovery path was, what alternative exists, and what evidence shows the repair lasted. The organization needs to map dependencies, test recovery, report risk in terms of business functions, and publish status without confusing communication with performance.

ARIN’s record demonstrates both progress and a boundary. It shows a candid account, separate recovery times, ordered follow-up, recurring technical-debt oversight, a service taxonomy, and later management-reported mitigation. It does not show every hidden dependency, a public recovery-time objective for the affected services, or an independent audit of every completed control. That is not an indictment and not a clean bill of health. It is an evidence-led conclusion: the path from redundancy to resilience is visible only when the service, the failure, the recovery, and the proof are kept separate.

Sources

  1. ARIN Board of Trustees and John Curran biography
  2. ARIN Organization Structure & Staff
  3. Board of Trustees Meeting Minutes — 16 January 2019
  4. Richard Jimmerson, “Operations at ARIN: New Blog Series and Recent Outage Information,” 19 March 2020
  5. Board of Trustees Meeting Minutes — 22–23 January 2020
  6. Board of Trustees Meeting Minutes — 25 March 2020
  7. Board of Trustees Meeting Minutes — 21 May 2020
  8. Meeting of the ARIN Board of Trustees — 3 February 2021
  9. ARIN, “New ARIN Service Status Page Available,” 5 April 2021
  10. Meeting of the ARIN Board of Trustees — 4 August 2022
  11. ARIN 2025 Annual Report
  12. Visual identity reference only: ARIN’s official John Curran photograph. Used only to ground the AI editorial portrait; it is not evidence for the operational claims.