Summary

  • Ofcom reported that Shetland residents lost mobile and fixed-line services for approximately 16 hours in October 2022 after the south subsea fibre link was damaged while a secondary link was already unavailable because of damage sustained the previous week. [10]
  • Ofcom said legacy low-bandwidth microwave backhaul kept PSTN voice largely available. It also said temporary restoration was achieved by increasing optical power at the source and that repairs across both primary and secondary links took about ten days. [10]
  • Scottish Government correspondence records intermittent problems on the resilient link from 14 October and describes the primary link as damaged accidentally by a UK-registered fishing vessel shortly after midnight on 20 October. The public record does not support a sabotage claim. [4][5][11]
  • Faroese Telecom statements reported contemporaneously said technicians activated spare or unused fibre capacity in the damaged cable and restored much of the service carried on BT fibre. That is evidence about a specific recovery action, not proof that every service or provider recovered at the same time. [15][20]
  • Shetland Islands Council said some customers remained offline because their providers did not have resilient connections and could not access undamaged fibre. The Council also reported problems affecting some telecare and community-alarm services. [13]
  • Police Scotland deployed officers and vehicles to preserve access to emergency services while normal communications were impaired. [12]
  • The accountability question is not whether two physical cables existed. It is who controlled usable alternate capacity, provider access, optical and microwave fallback, emergency arrangements, repair sequencing and the evidence used to declare services restored.
  • Complete provider contracts, routing policies, optical budgets, customer populations and loss records are not public. The evidence supports a control analysis, not a provider-by-provider accusation.
  • The nearest comparison is Tonga's 2022 cable break, but the cases are distinct. Tonga centered a sole international route, volcanic seabed damage and long repair logistics. Shetland exposed the gap between nominally diverse routes and service-specific access to surviving capacity.

A major outage can exist inside a redundant diagram

Redundancy is one of the most comforting words in infrastructure.

It suggests that the failure of one component will not become the failure of a service. A second cable, router, power feed, data centre or control plane is expected to carry the load. The word often appears in architecture diagrams, procurement documents, regulator submissions and executive risk registers long before anyone sees whether it works.

Shetland's October 2022 outage shows why installed redundancy and usable redundancy must be treated as different facts.

The islands were connected through subsea fibre routes running south through Orkney toward mainland Scotland and north toward the Faroe Islands. The public record describes one as the primary route and the other as a resilient or secondary route. In the week before 20 October, the northern link was already impaired. Shortly after midnight on 20 October, the southern link was damaged. Ofcom later summarized the result as an approximately 16-hour loss of mobile and fixed-line services. [4][10]

At first glance, the event can be described as bad luck: two physical faults within days. The damage was reportedly accidental. A fishing vessel was identified in the public record. Cable repairs require ships, weather windows, specialist crews and physical access to fault locations. None of those facts is trivial. [4][11][16]

But bad luck does not answer the operational question.

Why did a system with two routes produce such wide disruption? Why did some services return quickly while some customers remained affected for much of the following week? Why could legacy voice remain more available than data-dependent services? Why could one set of traffic use restored or alternate capacity while another could not?

Those questions move the analysis away from the existence of cable steel, glass and landing equipment. They move it toward the controls that convert infrastructure into service:

  • route availability;
  • optical power and margin;
  • usable fibre pairs;
  • reserved and spare capacity;
  • switching and routing configuration;
  • wholesale access;
  • retail-provider arrangements;
  • fallback backhaul;
  • emergency-service plans;
  • repair priority;
  • service monitoring;
  • restoration evidence.

A second path can be physically present but unavailable because it is already damaged. It can be healthy but lack enough capacity. Capacity can exist without a contractual or technical path for a particular service. A provider can have a route on paper but fail to test the failover state. A low-bandwidth backup can preserve voice while broadband remains unavailable. An operator can restore light through fibres without proving every downstream service has recovered.

The Shetland record contains examples of several of those conditions. That is what makes it a network-infrastructure accountability case rather than a generic story about a cable cut.

The two faults must remain separate

The event needs two clocks.

Scottish Government correspondence says the resilient cable had experienced intermittent issues since 14 October and that Faroese Telecom had dispatched people to work on it. The second event occurred shortly after midnight on 20 October, when the primary cable was damaged. The correspondence describes the cause as accidental damage by a UK-registered fishing vessel and says the event affected almost all services carried on that cable. [4]

Ofcom's later account uses slightly different operational language. It says a secondary fibre link was already unavailable because of damage sustained the previous week. It says the south fibre cable then suffered a loss of light. [10]

Those descriptions should not be flattened into a claim that two cables were cleanly severed at the same moment.

The public record distinguishes intermittent trouble, unavailability, fibre effects, loss of light and cable repair. Faroese Telecom's managing director was reported as saying that the later break affected fibres but did not cut off the entire cable. Local reporting places the fault a few kilometres off Shetland and describes a repair ship working at more than one location. [15][16][18]

That distinction matters for three reasons.

First, the physical mechanism affects the recovery options. A total break, a damaged fibre pair, an amplifier problem, a power-feed problem and a reduced optical margin are not the same fault. They may allow different temporary measures.

Second, the description affects responsibility. A cable owner may control repair and optical restoration. A carrier may control traffic placement. A retail provider may depend on a wholesale path. A public authority may control emergency response. Assigning one label to the whole event can hide those boundaries.

Third, the distinction prevents retrospective exaggeration. The incident was serious without calling every cable completely severed or every service totally unavailable for ten days. Ofcom's approximately 16-hour figure describes the main loss of mobile and fixed-line services. The ten-day period describes work across both fibre links before permanent restoration. Some residual service problems continued between those points. [10][13][18]

An accountable chronology therefore separates at least five states:

  1. The northern or secondary route develops problems.
  2. Repair work begins while the southern route still carries service.
  3. The southern route is damaged and broad service loss follows.
  4. Temporary optical and capacity measures restore much of the traffic.
  5. Residual provider-specific faults continue until permanent repairs complete.

That sequence is more useful than a single outage duration because each state exposes different controls and decision owners.

Physical diversity is not service diversity

Network resilience is often measured at the wrong layer.

At the physical layer, Shetland had more than one direction for fibre connectivity. One route ran toward Orkney and the Scottish mainland. Another ran toward the Faroe Islands. The geography appears diverse.

At the service layer, however, a customer does not buy an abstract direction across the seabed. The customer buys a broadband, mobile, voice or institutional service. That service travels through a chain of access equipment, local aggregation, wholesale transport, provider routing, authentication, core systems and upstream interconnection. Every link in that chain needs a working alternate if the service is to survive a cable fault.

The Council's 28 October update is unusually important because it states the service-layer problem directly. It said users who remained without internet were served by companies that did not have resilient connections. Those providers could not access the undamaged fibre and were dependent on repairs. [13]

The statement does not publish every contract or routing table. It does not prove why each provider lacked access. It does not show whether the limitation was commercial, technical, historical or some combination. It should not be converted into a broad claim that named companies deliberately refused resilience.

What it does prove is that alternate physical capacity did not translate automatically into alternate service for all customers.

This is a recurring infrastructure failure pattern.

A data centre can have two power feeds that share an upstream substation. A cloud application can run in two availability zones while depending on one identity control plane. A mobile operator can have redundant transport while both routes terminate on the same router. A registry can have multiple name servers that receive the same malformed zone. An island can have two cable directions while a provider's service is reachable through only one operational or commercial path.

In each case, counting components overstates resilience. The correct question is whether failure domains are independent all the way to the service.

For Shetland, that means testing:

  • whether the north and south paths were physically independent at the fault locations;
  • whether landing and terrestrial segments avoided common points;
  • whether sufficient capacity existed on the alternate path;
  • whether providers were authorized and configured to use it;
  • whether failover could occur automatically;
  • whether manual rerouting procedures were current;
  • whether service dependencies such as DNS, authentication and management remained reachable;
  • whether emergency and public-service traffic had priority;
  • whether customers could be measured by restored service, not only restored light.

The public record answers some of those questions and leaves others open. Those evidence gaps are themselves part of accountability and should remain explicit.

Increased optical power was a recovery control

Ofcom's account contains a technical detail that changes the story.

It says the south fibre cable experienced a loss of light and that temporary restoration was achieved by increasing power output at the source. [10]

Optical fibre systems depend on a link budget. The signal launched into the fibre loses power as it travels through the cable, repeaters, connectors and other components. A receiver needs enough signal to distinguish data reliably. Damage can increase loss without destroying every fibre. If the remaining path has enough margin, changing the launch conditions may restore a usable signal temporarily.

The public report does not publish the exact optical design, power levels, fibre pairs or engineering thresholds. It would be wrong to reconstruct those details from general principles. It is still reasonable to identify the control surface.

Someone had to:

  • detect that the problem was a loss-of-light condition;
  • determine which fibres or services remained recoverable;
  • assess whether additional power was safe and useful;
  • authorize the change;
  • monitor errors and stability;
  • decide what traffic could return;
  • keep permanent repair work moving.

That is not the same control as owning a second route.

Optical restoration also creates an evidence requirement. "Light restored" is a necessary network fact, but it is not the same as "customer service restored." A fibre can pass a signal while routing, authentication, capacity or downstream provider paths remain impaired. The recovery record must therefore bind physical measurements to service tests.

An accountable restoration report would separate:

  1. optical continuity;
  2. link-layer stability;
  3. routed reachability;
  4. provider handoff;
  5. broadband session establishment;
  6. mobile backhaul and core reachability;
  7. voice service;
  8. emergency calling;
  9. public-service application access;
  10. customer-observed availability.

The Shetland incident shows why this layered approach matters. Much of the service returned on 20 October, yet the Council was still warning about internet and telecare problems on 28 October. [13]

Temporary engineering success and complete service recovery were not the same milestone.

Spare fibre mattered because it could be activated

Contemporaneous local reporting adds another recovery mechanism.

Shetland News reported that Faroese Telecom activated spare and unused fibre capacity in the damaged cable. The report said this capacity had been reserved for later use and that activating it helped reconnect services carried on BT fibre. The operator's managing director said the fault affected fibres without cutting off the whole cable and that technicians restored services to the extent possible. [15]

The wording is attributed reporting, not an Ofcom engineering annex. It should be used with that level of confidence. Even so, it demonstrates an important principle: spare infrastructure has value only if responders can activate it during the incident.

Spare fibre can exist in several forms. A cable may contain unlit fibres, unused capacity on lit pairs or capacity reserved for growth. Turning that reserve into service can require equipment, configuration, rights of use, compatible interfaces and operational coordination.

The accountability questions are practical:

  • Was the spare capacity inventoried?
  • Did operators know its condition?
  • Was it connected to equipment that could be used during a fault?
  • Were activation procedures tested?
  • Who had authority to allocate it?
  • Which services had priority?
  • Could downstream providers reach it?
  • Was there enough margin for sustained operation?
  • What monitoring showed that the temporary path was stable?

The incident suggests that at least some reserved capacity was convertible into real restoration. It also suggests that the conversion did not cover every downstream service.

This is a more useful lesson than saying "build more cables." More physical assets may reduce risk, but a cable that cannot be reached, switched, powered, contracted or operated during a fault is an expensive diagram. Resilience investment needs operational acceptance tests.

For each alternate path, a provider should be able to show:

  • the services mapped to the path;
  • the capacity available under failure;
  • the trigger for failover;
  • the commands or automation used;
  • the people authorized to act;
  • the dependencies that remain outside the path;
  • the test results;
  • the recovery objective;
  • the residual risk.

Without that evidence, spare capacity is a possibility rather than a control.

Low-bandwidth microwave preserved a different kind of continuity

Ofcom says PSTN landlines for voice remained largely available because of legacy low-bandwidth microwave backhaul. [10]

That sentence should not be read as proof that all landlines worked or that emergency communication was unaffected. Scottish Government situation reporting described broad loss across landline, internet and mobile services, and Police Scotland still deployed a physical presence to preserve emergency access. [4][12]

The two records can coexist.

Legacy microwave may have preserved a bounded voice path while high-bandwidth services failed. Availability can vary by exchange, provider, service, location and time. A low-bandwidth link that is limited public evidence for normal internet traffic may still carry some voice. Conversely, a customer relying on a digital or provider-specific service may not benefit from the legacy path.

This is a reminder that fallback should be designed around minimum viable functions, not only normal capacity.

During a severe communications failure, priorities may include:

  • emergency calls;
  • coordination among police, fire, ambulance and coastguard;
  • hospital and health-board communications;
  • telecare alarms;
  • public warnings;
  • essential government contact;
  • critical payment and logistics functions;
  • repair-team management.

A fallback path can be valuable even if it cannot reproduce ordinary broadband. The engineering and governance task is to identify which services it must support and prove that they can use it.

The Shetland incident also shows the danger of treating old infrastructure as accidental resilience. If a legacy microwave path is carrying a safety function, its role should be documented, maintained and tested. If it is expected to be retired, the replacement architecture needs an equivalent or better failure path before retirement.

Otherwise, modernization can remove a hidden layer of continuity.

Legacy systems should not be romanticized. Low bandwidth, limited coverage and ageing equipment create their own risks. The accountable point is narrower: a fallback's value is determined by the service it can preserve, not by whether it is old or new.

Emergency access turned network recovery into public duty

Communications outages become public-safety events when people cannot reliably reach emergency organizations.

Police Scotland said officers and vehicles were moved into the area to ensure the public had access to emergency services. After communications returned to normal, those resources were withdrawn. [12]

Contemporaneous reporting described advice to try 999 from a landline or mobile and, if that failed, to flag down an emergency vehicle or attend a police, hospital, fire or ambulance location. Extra patrols and a physical emergency hub were used as substitutes for normal reachability. [19]

These workarounds matter because they reveal the true service being protected.

An emergency number is not just a telephone feature. It depends on:

  • access-network availability;
  • call routing;
  • caller location where technically available;
  • interconnection with emergency organizations;
  • power at customer and network sites;
  • public knowledge of alternatives;
  • staffing and physical reachability when communications fail.

Ofcom's General Conditions describe requirements aimed at broad availability, uninterrupted access to emergency organizations and arrangements for provision or rapid restoration during a disaster. The exact regulatory duties applicable to a provider and event require legal analysis, and current consolidated guidance should not be applied retroactively without care. It nevertheless shows why emergency access is treated as more than an ordinary customer-experience metric. [8]

The Shetland response should be evaluated at two layers.

The first is network prevention and recovery. Were emergency-call paths diverse? Did the remaining voice backhaul work? Could mobile calls reach the emergency network? Did providers have tested arrangements with emergency organizations?

The second is civil contingency. When network reachability became uncertain, did authorities establish visible alternatives, move resources, check vulnerable people and communicate through channels that still worked?

Police and local authorities controlled the second layer. They did not control cable condition or wholesale routing. Their response reduced harm, but it should not be used to excuse weaknesses in the network layer.

A good emergency workaround is evidence of resilience. It is not evidence that the original service met its availability objective.

Telecare exposed the dependency below the application

Shetland Islands Council reported problems affecting some telecare and community alarms, including services delivered over analogue landlines. Staff contacted clients and asked family and friends to check elderly or vulnerable relatives. [13]

Telecare is often discussed as a device or social-care application. During an outage, its network dependencies become visible.

A pendant or home alarm may appear simple to the user. Behind it sits a chain:

  • power at the premises;
  • the alarm unit;
  • local wiring or radio;
  • a landline, mobile or internet access path;
  • exchange and backhaul;
  • a monitoring centre;
  • a human response process.

The device can be functional while the path is unavailable. A network can be partly restored while a specific alarm still fails. A provider can report broad service recovery while a vulnerable user remains isolated.

This is why service restoration must be measured by critical function, not only aggregate traffic.

The Council's update does not say every alarm failed. It says some services were affected and describes a manual welfare response. That boundary is important because a risk signal must not be inflated into an unsupported injury claim.

The supported conclusion is that telecom continuity is part of care continuity. Operators and public agencies need a shared dependency map that identifies:

  • which alarm types use which access technologies;
  • which provider and wholesale paths they depend on;
  • what battery or power assumptions apply;
  • how failure is detected;
  • who receives an alert when the communication path fails;
  • what manual welfare process begins;
  • how restoration is verified at the device.

Shetland's experience turned an abstract network outage into a specific question: could a vulnerable person still call for help?

That question belongs in infrastructure design before the cable breaks.

Provider-specific outcomes are evidence, not a licence to speculate

Residual service differences are among the most revealing parts of the event.

The Council said some affected users were served by companies without resilient connections and could not access the undamaged fibre. Local reporting later described some providers as rerouting while others remained unavailable. [13][18]

It is tempting to turn that record into a league table of responsible and irresponsible providers. The evidence does not support such a simple conclusion.

The public sources do not disclose:

  • every wholesale contract;
  • every route purchased by each retail provider;
  • capacity limits on alternate paths;
  • the exact state of each provider's equipment;
  • which customers were on which infrastructure;
  • whether a failover attempt failed;
  • whether a route was technically possible but commercially unavailable;
  • which service-level objectives applied;
  • what information providers received during repair.

A provider may lack an alternate because it did not purchase one. It may depend on a wholesaler that lacks one. An alternate may exist but have limited public evidence capacity. The service may share a local failure domain that survives neither direction. A manual process may be too slow. Public reporting cannot resolve all of those possibilities.

The accountable approach is to ask for evidence.

Each provider that sold service into a geographically constrained market should be able to state, at an appropriate level of security:

  1. the failure domains in its access and backhaul path;
  2. the alternate route available when one subsea segment fails;
  3. the capacity committed on that route;
  4. the conditions that trigger failover;
  5. the most recent end-to-end test;
  6. the services that will degrade or remain unavailable;
  7. the recovery target;
  8. the customer and public-authority communication plan.

That is not a demand to publish sensitive topology. It is a demand to prove that "resilient" is attached to an operational control.

The Council's statement is valuable precisely because it reports observed service behavior. Some users could not use undamaged fibre. That fact should drive a focused inquiry into why, without assuming the answer.

Repair sequencing allocated continuity

Physical cable repair is a constrained operation.

The fault must be located. A suitable vessel and specialist crew must be available. Weather and sea conditions must permit work. The cable may need to be recovered, cut, spliced, tested and returned to the seabed. If more than one section is damaged, the vessel cannot necessarily repair both at once.

Local reporting said the repair ship completed work on another section west of Shetland before moving to the fault east of the islands. Full broadband restoration was reported after the later repair was completed on 30 October. [18]

This sequence means repair order is also an allocation decision.

Operators need criteria for deciding:

  • which fault creates the greatest service risk;
  • which repair is physically possible first;
  • whether a temporary path is stable enough to wait;
  • which route restores the most critical services;
  • whether another fault would create total isolation;
  • how long backup capacity can carry the load;
  • whether equipment and spares are available;
  • what weather window is expected.

The public record does not show the complete decision log. It does show that repairs across both primary and secondary links took about ten days and that temporary controls were used during the interval. [10]

Accountability does not require pretending that subsea repair can be instantaneous. It requires evidence that the operator understood the changing risk and made the best controllable choices.

That evidence might include:

  • fault-location confidence;
  • service-impact estimates;
  • vessel and crew availability;
  • route-capacity state;
  • weather constraints;
  • temporary restoration stability;
  • emergency-service requirements;
  • repair-order rationale;
  • updated completion estimates;
  • post-repair optical and service tests.

Uncertainty should be visible. A forecast based on vessel arrival and weather is not a promise. A public status update should distinguish known fault state, current workaround, estimated repair time and remaining risk.

The Shetland event shows why repair readiness is part of network design. A route is not fully resilient if the operator cannot locate faults, obtain a vessel, access spares, coordinate landing sites and communicate realistic recovery evidence.

Practical control was divided

The event involved several organizations, but control was not evenly distributed.

The cable owner and operator

Faroese Telecom and the SHEFA system sat closest to physical fault detection, optical condition, spare fibre and permanent repair. Operator-attributed statements describe temporary technical actions and cable-ship work. That gives the cable operator significant control over physical recovery.

It does not give the operator control over every retail customer's service design or every downstream contract.

Wholesale and retail communications providers

Providers controlled, or depended on others to control, traffic placement, alternate-route access, capacity, customer communication and service restoration. The Council's report of provider-specific continuity makes this layer central.

The public record does not reveal which part of every provider chain made the decisive choice. Responsibility should follow verified capability, not brand visibility.

Ofcom and government

Ofcom controlled regulatory reporting and the broader resilience framework, not incident repair. Its Connected Nations account provides a public technical baseline. Government resilience teams coordinated information and emergency response. The regulatory and policy question is whether providers were required to understand and reduce foreseeable common failure.

The strengthened telecom security and resilience framework came into force in October 2022. Ofcom described duties to identify risks, reduce the effect of security compromises and take remedial action. A physical cable accident is not automatically a security compromise within every legal provision, so the framework should be used as resilience context rather than an unsupported enforcement finding. [6][7]

Police, Council, health and emergency partners

Public authorities controlled local mitigation: patrols, physical access points, welfare contact and service coordination. They absorbed consequences created by the communications failure but did not control the cable or provider routing.

Their work is evidence of public cost transfer. When normal telecommunications fail, public organizations supply people, vehicles and manual checks to recreate part of the service.

Customers and vulnerable users

Customers held evidence of failed service and personal consequences. They generally did not control route diversity, repair or alternate capacity. Advice to use another device, move location or check a neighbour may reduce harm, but it does not transfer infrastructure responsibility to the user.

This control map avoids two errors.

The first is blaming the actor nearest the physical damage for every downstream consequence. The second is treating distributed responsibility as no responsibility. Different parties controlled different prevention, continuity, recovery and evidence functions.

Accountability does not require proving malicious intent

The public record describes accidental damage by fishing activity. [4][11][16]

That finding narrows the event. The available evidence does not support framing the outage as sabotage, geopolitical attack or deliberate disruption.

It does not remove accountability.

Infrastructure is often harmed by ordinary causes: excavation, anchors, trawling gear, weather, failed maintenance, power loss and configuration mistakes. A resilience system exists because those causes are foreseeable even when their exact timing is not.

The accountability questions are therefore:

  • Was the route protected and monitored appropriately?
  • Was the second route genuinely independent?
  • Did services have usable access to it?
  • Was failover tested?
  • Was backup capacity sufficient?
  • Could responders restore degraded fibres?
  • Were emergency arrangements current?
  • Were repairs ready?
  • Was the public told what remained unavailable?
  • Did post-incident work reduce recurrence?

None of those questions depends on malicious intent.

Intent matters for legal and security attribution. Control matters for operational accountability.

The same distinction protects against unfair blame. A cable owner cannot prevent every fishing vessel from contacting a cable. A retail provider cannot repair an undersea fibre it does not own. A local Council cannot reconfigure a carrier's backhaul. The fair inquiry asks what each party could realistically do before, during and after the event.

An operator may be accountable for route marking, monitoring, fault localization and repair readiness. A carrier may be accountable for alternate capacity and tested failover. A regulator may be accountable for resilience expectations and evidence. Public agencies may be accountable for emergency plans.

This is stronger than a blame narrative because it identifies controls that can change.

Impact figures need layer discipline

The incident produced several different time and impact measures.

Ofcom's approximately 16 hours refers to the main period in which residents were cut off from mobile and fixed-line services. The regulator also describes ten days of repair work across both links before permanent restoration. [10]

Local sources report that some customers remained disrupted during that interval, particularly where providers could not use alternate capacity. [13][18]

Those figures should not be combined into one claim that all customers were offline for ten days.

Likewise, reports of a "complete outage" should be read in context. Government situation reporting said most landline services, all internet services and all mobile services were affected at a particular point. Ofcom later said PSTN voice remained largely available through microwave backhaul. [4][10]

These records may reflect different timestamps, service definitions and information available to each institution. The responsible approach is to present both and avoid inventing a false precision.

A useful accounting distinguishes:

  • broad service-loss period;
  • temporary restoration time;
  • residual provider-specific disruption;
  • permanent physical repair;
  • individual customer downtime;
  • emergency workaround duration.

Any restoration account must also separate users, subscriptions, premises and services. One household can have fixed broadband, multiple mobile subscriptions and a landline. A telecare alarm can fail while another voice service works. A provider can restore aggregate traffic while a specific group remains disconnected.

Without provider-level records, a complete affected-user count or financial loss cannot be produced.

That limitation is not a weakness to hide. It is a finding about public evidence quality.

Restoration claims need a matrix, not one timestamp

The phrase "services restored" is necessary for public communication, but it can conceal uneven recovery.

For Shetland, at least four restoration statements were true at different times:

  1. A temporary technical solution restored much of the broad service on 20 October.
  2. Police later said normal communications had returned sufficiently to withdraw extra resources. [12]
  3. Some customers and telecare functions still had problems on 28 October. [13]
  4. Full broadband restoration was reported after permanent repair on 30 October. [18]

A restoration matrix would make those differences explicit.

Rows should represent service functions:

  • PSTN voice;
  • mobile voice;
  • emergency calling;
  • mobile data;
  • fixed broadband;
  • business connectivity;
  • Council services;
  • health systems;
  • telecare;
  • payment systems;
  • airport and transport communications.

Columns should represent evidence:

  • optical link state;
  • route reachability;
  • session establishment;
  • successful transactions;
  • external probes;
  • provider telemetry;
  • public-authority confirmation;
  • customer observations;
  • residual incident count.

A service should move from "degraded" to "restored" only when the required evidence is present. The standard need not be identical for every service. Emergency calling may require a stronger test than ordinary web browsing. A provider may need to state that service is restored for most users while a known cohort remains affected.

This avoids two harmful extremes.

One is waiting for perfect certainty before communicating any progress. The other is declaring total restoration when only the core path has returned.

Shetland's layered recovery is exactly the kind of event that benefits from a matrix.

The regulatory lesson is evidence, not slogans

Ofcom included the Shetland event in its Connected Nations 2022 discussion of network and service resilience. That publication gives the incident a regulatory significance beyond local reporting. [9][10]

The lesson should not be reduced to "more regulation" or "more investment."

An effective resilience framework needs evidence that providers:

  • identify critical dependencies;
  • understand geographic and logical common modes;
  • maintain sufficient alternate capacity;
  • test failover under realistic load;
  • protect emergency access;
  • monitor degradation;
  • communicate major incidents;
  • preserve repair capability;
  • learn from observed failure.

Ofcom's current General Conditions describe broad availability and emergency-planning expectations. Its network-security work describes a framework for provider risk reduction and remediation. The precise legal application depends on provider, service and date, and no event-specific compliance finding follows from the available public evidence. [6][7][8]

What the event supports is a verification principle.

Regulators should not accept the existence of a second route as proof of resilience. They should ask which services can use it, at what capacity, under what trigger, and with what test evidence.

This is running-code primacy applied to telecom continuity. The diagram is a claim. Observed failover is evidence.

The distinction also improves investment decisions. If the failure was caused by lack of physical diversity, build another route. If alternate capacity existed but providers could not reach it, fix access and interconnection. If failover was too slow, improve automation and procedures. If emergency systems depended on one access technology, diversify the service path. If repair logistics dominated recovery, improve spares and vessel arrangements.

Different failure mechanisms require different remedies.

Shetland is not Tonga

A useful comparison is Tonga's 2022 cable break. Both events involve islands, subsea fibre, fallback capacity, repair ships and public-service dependence.

The common vocabulary does not make them duplicates.

Tonga's principal international cable was severed after the Hunga Tonga-Hunga Ha'apai eruption and tsunami. The event involved volcanic seabed processes, a sole international fibre route, constrained satellite fallback, international repair logistics and a much longer restoration story for domestic connectivity.

Shetland's October 2022 event involved two nominally diverse directions impaired within a week. The record shows temporary optical restoration, activation of spare fibre, legacy microwave voice fallback and different outcomes among providers that could or could not use alternate capacity.

The control questions are different.

For Tonga:

  • Why was there only one international fibre route?
  • What fallback capacity existed?
  • Were spares and repair vessels ready?
  • How should a country manage a long repair after an extreme natural event?

For Shetland:

  • Were two installed routes operationally independent?
  • Which services could use the remaining or restored capacity?
  • Did providers have sufficient alternate access?
  • How did optical, spare-fibre and microwave controls work?
  • Why did residual outages differ by provider?

The comparison is useful only if it sharpens those boundaries.

What the public record cannot prove

Several material facts remain unknown.

The public record does not reveal:

  • full SHEFA-2 fibre and equipment topology;
  • exact optical budgets before and after damage;
  • all route configurations;
  • traffic volumes by service;
  • capacity reservations;
  • every provider's wholesale agreements;
  • precise provider-by-provider customer counts;
  • complete failover logs;
  • internal incident commands;
  • repair prioritization records;
  • full financial losses;
  • all telecare failures;
  • independently tested remediation.

It also does not establish that the northern and southern incidents had one cause, that either was malicious, or that every named institution failed a legal duty.

These gaps matter because they limit the allocation of responsibility.

They also define what a stronger public post-incident record should contain:

  1. a bounded topology showing failure domains without exposing sensitive detail;
  2. a service-impact timeline;
  3. the alternate capacity available;
  4. failover results by service class;
  5. emergency-access evidence;
  6. temporary restoration actions;
  7. repair sequencing and constraints;
  8. residual outage cohorts;
  9. root-cause confidence;
  10. remediation tests.

A high-quality post-incident report does not need to reveal exploitable network detail. It needs enough evidence for customers, public authorities and regulators to distinguish unavoidable physical damage from controllable continuity failure.

A usable-redundancy accountability test

The Shetland event supports a reusable test for island and remote-region connectivity.

1. Map the service, not only the asset

Identify every path from customer access through local aggregation, subsea transport, provider core, authentication and external interconnection. Mark common landing, power, routing and management dependencies.

2. Prove geographic independence

Document physical separation at likely damage zones, landing points and terrestrial segments. A north and south label does not prove that all relevant segments are independent.

3. Reserve usable capacity

State how much alternate capacity is committed during failure. Include critical services and degraded operating modes.

4. Bind providers to the alternate

Show that each dependent provider has the technical and commercial ability to use the route. A healthy fibre that a service cannot access is not redundancy for that service.

5. Test end-to-end failover

Run controlled exercises that verify routing, sessions, DNS, authentication, voice, emergency calling and public-service applications. Record the result and repair the gaps.

6. Define minimum viable continuity

Specify which functions must survive when normal broadband capacity cannot. Voice, emergency coordination and telecare may need distinct fallback paths.

7. Inventory dormant controls

Spare fibres, microwave links and satellite terminals should have owners, activation procedures, compatibility tests and monitoring.

8. Prepare physical repair

Maintain fault-location capability, spares, cable-ship arrangements, landing access, weather assumptions and repair priorities.

9. Measure restoration by service

Use a matrix that distinguishes optical restoration, route restoration and customer-service restoration. Preserve residual cohort data.

10. Publish bounded evidence

Explain what failed, what worked, what remained unknown and how remediation was tested. Do not use the word "resilient" as a substitute for evidence.

Conclusion

The 2022 Shetland outage was triggered by two physical cable problems within a short interval. That timing was unlucky. The distribution and duration of service harm were shaped by controls.

Ofcom recorded an approximately 16-hour broad loss of mobile and fixed services, legacy microwave support for much PSTN voice, temporary restoration through increased optical power and ten days of repair across both fibre links. Government and operator-attributed records describe accidental fishing-vessel damage, spare-fibre activation and precautionary satellite support. The Council documented provider-specific residual outages and telecare concerns. Police supplied physical emergency access while normal communications were uncertain. [4][10][12][13][15]

Together, those facts overturn a simple assumption.

Two cables do not equal two usable services.

Redundancy exists only when the alternate path has capacity, dependent providers can reach it, failover works, critical functions survive and restoration can be verified. A line on a network diagram is a design claim. The accountability evidence is what the network did when one line was already impaired and the other was damaged.

That is the lasting test from Shetland: not whether a second route was installed, but who could prove it was usable when continuity depended on it.

Sources

  1. https://news.sky.com/story/complete-outage-on-shetland-leaves-phones-internet-and-computers-unusable-12725347
  2. https://news.stv.tv/highlands-islands/internet-and-mobile-services-restored-in-shetland-after-power-outage-caused-by-damaged-subsea-cables
  3. https://news.stv.tv/highlands-islands/shetlands-damaged-subsea-cable-to-faroe-islands-now-repaired-confirms-operator
  4. https://www.gov.scot/binaries/content/documents/govscot/publications/foi-eir-release/2024/01-b/damage-to-the-telecommunication-cable-between-scottish-mainland-and-shetland-october-2022-foi-release/documents/202200331574---annex-a/202200331574---annex-a/govscot%3Adocument/202200331574%2B-%2BAnnex%2BA.pdf
  5. https://www.gov.scot/publications/damage-to-the-telecommunication-cable-between-scottish-mainland-and-shetland-october-2022-foi-release/
  6. https://www.ofcom.org.uk/internet-based-services/network-security
  7. https://www.ofcom.org.uk/internet-based-services/network-security/ofcom-begins-new-role-overseeing-security-of-telecoms-networks
  8. https://www.ofcom.org.uk/phones-and-broadband/accessibility/general-conditions-of-entitlement
  9. https://www.ofcom.org.uk/phones-and-broadband/coverage-and-speeds/connected-nations-2022
  10. https://www.ofcom.org.uk/siteassets/resources/documents/research-and-data/multi-sector/infrastructure-research/connected-nations-2022/connected-nations-uk-report.pdf?v=321647
  11. https://www.parliament.scot/chamber-and-committees/questions-and-answers?ref=S6W-12628
  12. https://www.scotland.police.uk/what-s-happening/news/2022/october/update-landline-and-mobile-phone-outage-shetland-resolved/
  13. https://www.shetland.gov.uk/news/article/2404/continuing-problems-affecting-connectivity-in-shetland-including-some-telecarecommunity-alarm-lines
  14. https://www.shetnews.co.uk/2022/10/21/cable-fault-the-third-reported-in-waters-around-shetland-since-mid-september/
  15. https://www.shetnews.co.uk/2022/10/24/perfect-storm-of-cable-failure-was-incredibly-bad-luck/
  16. https://www.shetnews.co.uk/2022/10/25/subsea-cable-fault-took-place-a-few-kilometres-off-shetland-telecoms-chief-says/
  17. https://www.shetnews.co.uk/2022/10/26/more-digital-disruptions-as-cable-repairs-get-under-way/
  18. https://www.shetnews.co.uk/2022/10/31/shetland-fully-connected-once-again/
  19. https://www.theguardian.com/uk-news/2022/oct/20/shetland-loses-telephone-internet-services-subsea-cable-damaged
  20. https://www.theguardian.com/uk-news/2022/oct/21/telephone-and-internet-restored-in-shetland-after-cable-damage