Summary
- BT's Public Emergency Call Service was disrupted from 06:24 to 16:56 UK time on 25 June 2023. The approximately 10.5-hour incident included an approximately one-hour nationwide total outage. Ofcom's final measure was 13,943 unsuccessful attempts from 12,392 unique callers, about 23% of attempts during the incident.
- The initiating configuration error was only the first part of the failure. Ofcom found that an incorrectly executed first failover and the reinstatement of the faulty node led to the total outage. The disaster-recovery platform then operated with important constraints, including a 50-call queue, degraded caller-location handling and no backup for relay calls.
- Ofcom found that BT contravened section 105A(1)(c) of the Communications Act 2003 and Regulation 9 of the Electronic Communications (Security Measures) Regulations 2022. It imposed a final £17.5 million penalty after a 30% settlement and admission discount. Those are the findings to report; the decision should not be enlarged to cover provisions Ofcom did not pursue.
- The public record did not confirm specific serious physical harm, but that leaves individual outcomes unresolved. Ofcom described substantial distress and a severe potential risk. The durable lesson is operational: emergency continuity depends on tested failover, sufficient recovery capacity, retained accessibility and location functions, and coordinated decisions across the whole call chain.
When the public consequence comes first
A failed retail call is inconvenient. A failed emergency call changes the risk facing a person who may already have very little time. The caller cannot inspect a network diagram, choose an alternative routing platform or determine whether a backup system has accepted the traffic. The service either creates a usable connection to an emergency authority or it does not.
That is why the 25 June incident should be read from the caller outward, not from the equipment inward. BT was the call-handling entry point for 999 and 112 calls and transferred them to emergency authorities. It did not operate every downstream emergency control room, and the wider response involved several organisations. Even so, the availability of BT's handling service was a necessary first step in the chain. If that step failed, the competence and capacity waiting downstream could not help a caller whose call had not reached them.
The scale was material. Ofcom recorded 13,943 unsuccessful attempts made by 12,392 unique callers over the incident. That represented approximately 23% of attempts. The figures do not mean that every unsuccessful attempt represented a separate emergency, or that every affected caller had the same outcome. They do show that failure was neither isolated nor momentary. Thousands of people met a service that did not perform its basic function when they tried to use it.
The timing deepens that concern. The disruption began at 06:24 and ended at 16:56, a window of approximately 10.5 hours. Within it was an approximately one-hour total outage. A long period of degraded service can create risk even when some calls still connect, because emergency systems are judged by the calls they fail as well as the calls they carry. A shorter interval of total loss is more acute, but the wider incident cannot be reduced to that hour. Recovery quality, not merely the moment at which some traffic returned, is part of the record.
The public record reviewed by Ofcom and the UK Government did not confirm specific serious physical harm. That is an important boundary, not a reassuring conclusion about every caller. Call outcomes were not comprehensively visible in the published material, and an absence of confirmation leaves the full pattern of outcomes unresolved. Ofcom found that the incident caused substantial distress and created a plausible severe risk. In an emergency network, exposure to that risk is itself a serious operational fact.
The service was a chain, not a single switch
An emergency call looks simple from the handset: a person dials three digits and expects an answer. Operationally, the call passes through a chain. BT's handling platform had to receive it, obtain or preserve information needed for handling, and transfer it to the relevant emergency authority. The receiving authority then had to answer and manage the incident. The government post-incident review treated the disruption as a whole-system problem because recovery depended on communication and coordination across those organisational boundaries.
This distinction matters for accountability. BT was responsible for the availability and resilience of the service it operated. Emergency authorities had responsibilities within their own operations. Government coordination and public communication had separate roles. It would be inaccurate to assign every consequence inside the wider emergency system to a single organisation. It would be equally inaccurate to treat the entry point as a minor component simply because other organisations completed the response. The first transfer is a control point: without it, the rest of the chain is unavailable to the caller.
The chain also explains why a backup must do more than accept an ordinary voice connection. Emergency handling depends on supporting functions. Caller-location information helps the receiving authority understand where assistance may be needed. Relay support enables certain callers with hearing or speech needs to use the service. Queue capacity determines how the platform behaves when demand arrives faster than it can be processed. If these features disappear or degrade during recovery, the service has not returned to an equivalent operating state.
For a non-specialist reader, “failover” means moving service from a primary system that cannot be trusted to a separate recovery system. “Disaster recovery” is the combination of technology and operating arrangements used to restore or sustain service after a serious failure. Neither term guarantees an outcome. A failover can be attempted incorrectly. A recovery platform can accept traffic yet lack capacity or functions. Procedures can exist but fail to guide a live decision. The practical question is always the same: what service did callers actually receive?
That question prevents a common reporting error. Describing a backup as “available” can sound like continuity was maintained. In this incident, the recovery environment was part of the eventual response, but its constraints mattered. The 50-call queue, degraded location handling and lack of relay-call backup were not secondary technical details. They defined what the recovered service could and could not do under the conditions for which it was needed.
A configuration fault became a continuity failure
The incident began with a technical and configuration fault. BT's public account described software and cache behaviour from the operator's perspective, while Ofcom's non-confidential decision provides the controlling regulatory account of the incident and its consequences. The important analytical point is not to collapse everything into the first fault. Complex systems experience faults. The resilience question is whether the controls surrounding that fault keep a critical service operating.
Here, the sequence after the initiating fault became central. Ofcom found that an incorrectly executed first failover and the reinstatement of the faulty node led to the total outage. In other words, the moment at which the organisation tried to move away from the impaired environment did not contain the problem. The way the transition was performed, followed by the return of the faulty component to service, intensified it.
That sequence turns a component-level explanation into an operating-control question. A configuration error can explain why a system first behaved incorrectly. It does not by itself explain why the organisation was unable to isolate the fault, choose the correct recovery state and move traffic safely. Those tasks depend on monitoring, decision criteria, documentation, training, access to accurate information and a shared understanding of who can authorise each action.
Ofcom's findings focused on those surrounding controls. Its decision addressed the adequacy of failover documentation, decision criteria, procedures, the training context and the disaster-recovery platform for a foreseeable risk. The regulatory conclusion therefore did not rest only on whether one piece of software failed. It examined whether BT had taken appropriate and proportionate measures to manage the security and resilience of a critical communications service.
This is the difference between a defect and a continuity failure. A defect is a condition in a system. A continuity failure occurs when the organisation's technical and human controls do not keep the required service operating through that condition. The public record supports the first description for the initiating event and the second for the wider outcome. Keeping them separate produces a more useful account than looking for a single dramatic cause.
It also avoids individual blame unsupported by the published evidence. The detailed record contains redactions, and it does not establish a complete map of individual decision ownership or vendor responsibility. The relevant public accountability lies in the design and operation of the service: whether procedures were usable, whether recovery capacity was adequate, whether alarms enabled diagnosis, and whether the service could be moved without reintroducing the fault.
The first failover was an operational event, not a diagram
Many continuity plans are easiest to understand when the system is healthy. A diagram shows a primary platform, a recovery platform and an arrow between them. The diagram answers where traffic is supposed to go. It does not answer how operators will recognise the correct moment to move it, what conditions they must check, how they will prevent the faulty state from following the traffic, or how they will know the recovery environment is coping.
The 25 June sequence exposed those missing dimensions. The first failover was incorrectly executed. The faulty node was reinstated. Those facts point to an operating transition that did not have sufficient protection against a wrong state. A robust transition would need clear entry criteria, a defined authority to act, steps that can be followed under pressure, a way to confirm that the source of failure has been isolated, and evidence that the destination platform is ready for the actual load.
Written instructions matter, but their existence is not the same as usability. A procedure can be technically complete yet difficult to apply during an incident. It can assume knowledge not held by the person on duty, rely on an ambiguous signal, omit a decision point, or fail to account for the interaction between systems. Testing matters for the same reason. It reveals whether the people, permissions, monitoring and platform behaviour align when the service must move, not just whether a document can be approved.
The phrase “tested failover” should therefore carry a demanding meaning for emergency services. It should include the operational sequence, not only the activation of backup equipment. It should show that traffic moves, that the faulty state remains isolated, that capacity is sufficient, that caller information survives, that accessible call paths remain usable, and that teams can reverse or stabilise the change without creating a second failure.
It should also be tested at the conditions that matter. A quiet demonstration may confirm that a recovery platform can accept a call. It may not show how the queue behaves under demand, how quickly an alarm becomes a correct diagnosis or how a receiving authority adjusts when features are degraded. Exercise design has to reflect the risk the control claims to manage.
This is why the article's central lesson is narrower than a slogan about redundancy. Redundancy is useful, but it is an input. Continuity is the observed ability to keep an essential service functioning through disruption. The distance between the two is filled by operational evidence: exercised transitions, measurable capacity, intact service features and decisions that work at incident speed.
Recovery restored a route, but with constrained service
After the total outage, the disaster-recovery platform became part of the route back to service. Ofcom's decision records limitations that are essential to understanding what “recovery” meant. The platform had a queue limited to 50 calls. Caller-location handling was degraded. Relay calls did not have a backup path. These limitations affected the ability of the recovery arrangement to reproduce the public service provided by the primary environment.
Queue capacity is not an abstract performance number. When more calls arrive than a system can immediately process, a queue holds them until capacity becomes available. A limit of 50 means the recovery platform had little room to absorb a sudden accumulation in a nationwide emergency service. Once that space was exhausted, additional callers could fail rather than wait. The very incident that requires a backup can also generate repeated attempts and concentrated demand, so capacity must be considered under stress rather than normal averages.
The second phase of the incident illustrates that pressure. Ofcom recorded 5,663 failed calls and an approximately 92% failure rate during that phase. A route existed, yet failure remained extremely high. That is a practical warning against defining recovery as the moment a backup is switched on. A service can be technically reachable and operationally inadequate at the same time.
Location handling matters because an emergency call is not just a voice session. The receiving authority may need reliable location data, particularly when a caller cannot give a clear address or the connection is interrupted. Degraded location capability increases the work required from the caller and the receiving team at the moment both can least afford it. The public record does not justify assigning a specific outcome to that degradation, but it establishes that recovery did not preserve the full normal function.
Relay calls make the same point from an accessibility perspective. A service is not fully resilient if its backup works only for callers who can use an ordinary voice path. The absence of relay-call backup meant that the recovery design did not provide an equivalent route for all users. Accessibility cannot be treated as an optional feature to restore after the core service; for the people who depend on it, it is the core service.
These constraints are related. A small queue can drive failures during a surge. Degraded information can make each successful connection harder to handle. Missing accessibility support can exclude some callers entirely. Continuity assurance must examine the combination, because a backup that loses several functions at once may impose a burden on both callers and emergency authorities even after traffic starts flowing again.
Four call counts, four different questions
The incident generated several prominent numbers. Used without definitions, they can look inconsistent. Used carefully, they describe different scopes, stages and purposes.
Ofcom's final measure was 13,943 unsuccessful attempts from 12,392 unique callers. The attempt count is higher because some people tried more than once. Ofcom also described unsuccessful attempts as approximately 23% of attempts during the incident. These are the controlling figures for the regulator's final quantified assessment.
The UK Government post-incident review used 9,641 unique callers in the context of callers requiring callback. That is a callback-scoped population, not a substitute for Ofcom's later count of all unique callers associated with unsuccessful attempts. A callback programme can apply eligibility or operational rules that produce a narrower set than the full incident measure.
BT's incident review reported a provisional figure of 11,470 unique unsuccessful callers. It was an earlier operator measure published before the final regulatory assessment. “Provisional” matters: organisations refine event records as duplicate attempts, time windows, system logs and definitions are reconciled. The existence of a later number does not make the earlier account deceptive, but the two should not be quoted as if they used the same final method.
The fourth important number is the 5,663 failed calls recorded by Ofcom during the second phase, with an approximately 92% failure rate. It is phase-specific. It helps explain the severity of a portion of the recovery period, but it is not the total for the whole 10.5-hour disruption.
Readers can keep the measures straight by asking four questions. Is the number counting attempts or people? Does it cover the whole incident or one phase? Is it a callback population or a service-failure population? Is it an operator's provisional figure or the regulator's final measure? If those labels travel with the number, the apparent conflict disappears.
The definitions also change what leaders should learn. Attempts help describe load and repeated efforts. Unique callers help describe breadth. A phase-specific failure rate shows how a system behaved during a particular operating state. A callback count helps manage a response obligation. No single measure answers every question.
This distinction is especially important in emergency-service reporting because repeated attempts are themselves operational behaviour. A person whose first call fails may immediately try again. That increases traffic at the same time the service has reduced capacity. A resilient design must therefore consider demand created by failure, not just ordinary demand. Metrics that separate attempts from callers make that feedback visible.
Ofcom's final finding and penalty
Ofcom closed its investigation with a final finding. It found that BT contravened section 105A(1)(c) of the Communications Act 2003 and Regulation 9 of the Electronic Communications (Security Measures) Regulations 2022. It imposed a £17.5 million penalty after applying a 30% settlement and admission discount.
The wording matters. This was not a proposed outcome. It was the regulator's final decision. At the same time, the finding should be reported only as far as Ofcom made it. The investigation considered a wider legal and technical context, but Ofcom did not pursue a finding under every provision examined. Expanding the holding would make the account less accurate, not more forceful.
The penalty provides a visible accountability outcome, but the decision is more valuable as a description of what resilience obligations require in practice. Ofcom examined measures around the operation and protection of the emergency-call service, including the foreseeability of the risk, the failover process and the capability of the recovery platform. Its conclusion connected legal duty to the functioning of the service, rather than treating compliance as a matter of possessing a plan.
That connection is important for other operators. A critical service can have modern equipment, specialist staff and a documented continuity design, yet still fall short if the controls do not work together in a foreseeable incident. The legal standard in this case was applied to the actual arrangement: the configuration risk, the decision path, the execution of failover, the recovery limitations and the resulting service impact.
The 30% discount should also be described precisely. It followed settlement and admission. The undiscounted mathematical starting point can be inferred, but it is not necessary to the public lesson and can distract from the final amount imposed. The reliable statement is that Ofcom imposed £17.5 million after the discount.
A fine cannot repair a failed call, and a regulatory decision does not operate a network. Its role is different. It establishes an authoritative record of the finding, identifies the control failures relevant to the duty and creates consequences for the regulated operator. Operational continuity still depends on the engineering, procedures, staffing and exercises that change what happens during the next fault.
Accountability belongs at the control boundaries
Incidents often invite a search for the person or component “at fault.” The published evidence here supports a more useful unit of analysis: the control boundary where a foreseeable problem should have been contained.
The first boundary was configuration and change control. A fault entered or existed in the live service and disrupted handling. The second was diagnosis: alarms and operating information had to help teams identify what was wrong. The third was the failover decision and execution. The fourth was isolation, including preventing the faulty node from being reinstated. The fifth was recovery capacity and functional equivalence. The sixth was coordination across BT, emergency authorities and government during a national public-service incident.
Each boundary asks a question that can be answered with evidence. What signal detected the failure? How quickly was the signal understood? Which written criteria authorised a move? Could the team execute the move without relying on memory? What check prevented an unhealthy component from returning? What traffic level had the backup carried in an exercise? Which normal features were preserved? Who told receiving organisations what had changed, and when?
These questions are harder than naming a broken component, but they are better predictors of future performance. Components change. Vendors change. Staff rotate. A well-designed operating control can survive those changes because it defines outcomes, ownership and evidence. A weak control stays weak even after one defect is patched.
The public record does not disclose every internal log or decision. That prevents a fair assignment of personal responsibility and makes speculation inappropriate. It does not prevent institutional accountability. Ofcom's final decision addressed BT's measures as the regulated provider. The government review addressed the wider system response. BT's own review recorded its account and immediate changes. Together, the sources show where assurance needed to become more concrete.
This framing also protects against an easy but incomplete remedy: fix the original software behaviour and declare the incident closed. Correcting the initiating fault is necessary. It does not prove that the next unrelated fault will be diagnosed, isolated and failed over correctly. The durable controls are those that work across fault types.
A national service needs whole-system coordination
The UK Government's post-incident review widened the lens beyond the BT platform. Emergency-call continuity depends on several organisations that do not share one control room or one chain of command. When the entry service is disrupted, receiving authorities need a common picture of what is failing, what routes remain available, what information callers may lack and what the public should do.
Coordination has technical and public dimensions. Technical teams need timely facts about service state and recovery changes. Emergency authorities need to understand likely call volumes and any loss of normal functions. Government needs to decide how to coordinate a national response. Public messages must be accurate enough to help without creating further confusion or avoidable load.
The callback effort shows how operational obligations continue after connectivity improves. A failed attempt can represent an unresolved request for help. Identifying and calling people back is not equivalent to preventing the failure, and the 9,641 figure used for that work is not the same as Ofcom's complete final unique-caller count. It is nevertheless part of the response: the system had to address people it had failed to connect.
This creates a second continuity requirement. The service needs not only a recovery path for new traffic but also a way to reconcile failed traffic. Logs must support the identification of affected callers within lawful and operational limits. Responsibility for callbacks must be understood. The receiving organisations need enough information to prioritise. Without that layer, the system can restore the front door while leaving earlier callers outside.
Exercises should therefore cross organisational boundaries. A platform test owned by one operator may confirm a technical switch but miss the questions that determine public outcomes. How will authorities be notified? How will degraded location information be handled? What happens to relay users? Who owns national messaging? How are unsuccessful attempts reconciled? These are not communications extras. They are part of the service delivered during disruption.
The government review's value lies in making that dependency visible. No single organisation can demonstrate whole-system continuity alone. Each can demonstrate its own controls, and the participants can jointly exercise the interfaces between them. The strength of the chain is revealed at those interfaces, especially when information is incomplete and decisions are time-sensitive.
Remediation is a claim that needs continuing evidence
The published record describes changes after the incident. Across Ofcom's decision, the government review and BT's account, the measures included improved alarms, clearer and tested failover arrangements, more automation, greater queue capacity, better caller-location handling, relay support and stronger coordination across the emergency-call system.
Those changes map sensibly to the failures observed. Better alarms can shorten the path from symptom to diagnosis. Clearer procedures and decision criteria can reduce ambiguity during a failover. Automation can remove error-prone manual steps, provided that automated behaviour is itself controlled and observable. More queue capacity can absorb bursts. Preserved location and relay functions can make recovery more equivalent to the primary service. Coordination can align the operator and receiving organisations.
The phrase “can” is deliberate. A completed action and an effective control are not always the same thing. Installing an alarm does not prove that the right team will interpret it correctly. Writing a procedure does not prove that it can be followed under pressure. Increasing a queue does not prove that capacity matches a credible demand surge. Automating a transition does not prove that it will isolate every faulty state.
The appropriate next question is what evidence now demonstrates the result. For alarms, evidence could include detection time in exercises and whether alerts identify the failing service rather than merely a downstream symptom. For failover, it could include end-to-end exercises, observed transfer times, isolation checks and restoration criteria. For capacity, it could include load tests above a defined emergency scenario. For accessibility and location, it could include successful test calls through the recovery path.
The evidence should be retained over time. A successful exercise immediately after an incident can lose relevance as software, network dependencies, staff and procedures change. Continuity is a maintained capability. Material changes should trigger new testing, and periodic exercises should cover different failure modes rather than repeating one rehearsed path.
Independent challenge also has value. Teams that build a control naturally know how it is intended to work. An exercise becomes more informative when participants face incomplete signals, competing hypotheses and an unexpected but bounded failure. The objective is not to surprise staff for its own sake. It is to test whether the system and its procedures guide sound action when the incident does not announce itself in the language of the plan.
None of the four sources used here provides a current, independent audit of every remedial control in 2026. It would therefore be wrong to claim that the published changes have permanently solved the risk. The fair conclusion is narrower: the reported measures address the failure modes identified, and their continuing effectiveness should be shown through current operational evidence.
What the public record cannot establish
The available sources are unusually detailed for a communications outage, but they are not complete. Parts of the regulatory decision are redacted. Internal logs, full architecture, vendor identity and individual decision records are not all public. Complete outcomes for every affected caller are also unavailable.
Those gaps impose limits. They prevent a precise reconstruction of every operator action and make it inappropriate to name a person or external supplier as the cause. They prevent a universal statement about caller harm. They also prevent readers from independently verifying the present condition of every remediation measure.
The gaps do not erase the findings that are public. Ofcom's final decision establishes the legal contraventions, the penalty and its assessment of the controls. The government review establishes a multi-party account of the operational response and lessons. BT's review provides the operator's account of the software behaviour, timeline and immediate actions. Where accounts use different figures, their definitions and dates explain how they should be used.
Disciplined reporting keeps facts at the level the evidence supports. “Ofcom found” is the right formulation for the legal conclusion and its account of the control failures. “BT said” is the right formulation for the operator's explanation of internal software and cache behaviour. “The government review reported” is the right formulation for the callback population and whole-system recommendations. Analysis can then draw lessons without turning inference into an attributed fact.
This discipline strengthens rather than weakens accountability. Overstatement gives an operator an easy way to challenge the story. Precision keeps attention on what is established: a national emergency-call disruption, a failed transition, a constrained recovery environment, thousands of unsuccessful attempts and a final regulatory finding.
The evidence leaders should ask for now
Boards and public authorities do not need to operate the platform, but they do need to ask questions that expose whether continuity is real. The first question is simple: when did the complete service last run on the recovery path under a representative load? A date, scope and result are more useful than a general assurance that tests occur.
The second question concerns equivalence. Which functions differ between primary and recovery environments? Queue capacity, location handling and relay support were material in 2023. A current assurance should identify whether those gaps are closed, how they are tested and what residual limitations remain. If a feature still degrades, leaders should know the operational consequence and the compensating control.
The third question concerns the transition. What exact condition triggers failover? Who has authority to act? Which steps are automated, and which require judgment? How does the system prove that the faulty component is isolated before service is restored? What prevents a well-intentioned action from reinstating the failed state?
The fourth concerns observability. An alarm is useful only if it reaches the right people with enough context to support a decision. Assurance should show time from fault to detection, detection to diagnosis, diagnosis to decision and decision to stable recovery. Those intervals reveal whether monitoring and procedures work together.
The fifth concerns demand created by failure. Load tests should include repeated attempts from callers who do not know whether the first call succeeded. They should test the queue at and beyond the credible peak, not only at an ordinary average. They should also examine how unsuccessful attempts are recorded and reconciled for possible callback.
The sixth concerns the wider system. When was the last exercise involving emergency authorities and government coordination, not only BT's technical team? Did it test notifications, public communication, location degradation, relay access and the transition back to normal service? Were actions assigned and retested?
The seventh concerns change. Which software, configuration, dependency or operating changes since the last exercise could invalidate the result? A recovery test is evidence about a particular system at a particular time. Configuration drift can quietly separate the tested design from the live one.
The final question is about residual risk. No critical system can promise that every fault will be prevented. Leaders should ask which failure modes remain capable of denying service, how quickly they can be detected and what independent route protects callers. An honest residual-risk statement is more valuable than a claim of complete resilience.
These questions translate a regulatory event into ongoing governance. They do not require a board to decide which server to restart. They require it to demand observable proof that the service can survive failure without losing the functions on which callers depend.
Continuity is an outcome, not an inventory item
BT's 2023 outage was not significant simply because a primary platform failed. It was significant because the surrounding controls did not preserve a national emergency service through that failure. The incorrectly executed first failover, reinstatement of the faulty node and constrained disaster-recovery platform turned an initiating configuration problem into a prolonged public-service disruption.
Ofcom's final decision gave that failure a legal consequence: findings under section 105A(1)(c) and Regulation 9, and a £17.5 million penalty after the 30% settlement and admission discount. The figures give it scale: approximately 10.5 hours of disruption, an approximately one-hour total outage, 13,943 unsuccessful attempts and 12,392 unique callers.
The deeper test is what happens next time. A backup platform, a procedure and an assurance report are useful only insofar as they produce a functioning service during a fault. For emergency communications, that means calls connect, demand is absorbed, location information is usable, relay users retain access, faulty states remain isolated and organisations coordinate the response. That is the evidence by which continuity should be judged.
Sources
Member Briefing
Deeper Profile Context
Sign in with the right membership level to unlock the full briefing and source notes.
Only for Strategic Circle
Strategic Circle
Open to all readers. Unlock profile briefings after joining and signing in.
Join Strategic CircleOnly for Leadership Alliance
Leadership Alliance
For qualified IP-asset owners and management; sign in to unlock alliance briefings.
Join Leadership Alliance
