Summary

  • The joint Dutch supervisory investigation places the principal KPN telephone outage between 15:34 and 18:52 on 24 June 2019. Fixed and mobile voice were almost entirely unavailable for KPN customers, and the ordinary path to 112 failed. Internet services continued to operate, so this was a call-routing and emergency-continuity failure rather than a total loss of connectivity. [1][2][5][7][10]
  • The failure crossed operator boundaries. The investigation found that 112 traffic from all fixed and mobile voice providers passed through KPN's telephone network before reaching the national public safety answering point. A platform operated by one company was therefore a shared national dependency. [3][5][7]
  • The technical mechanism defeated nominal redundancy. Four independently operating routing systems had counters that became synchronized following a software change. A warning script did not reset them as intended. The counters reached a negative state at almost the same time, error messages multiplied with repeated call attempts, and the platform stopped processing routing requests. [2][5][7]
  • A separate configuration problem impaired KPN's 4G delivery of NL-Alert messages. The joint report did not treat that fault as the cause of the voice outage. Its importance lies in simultaneous continuity pressure: a warning channel expected to help during the telephone failure was itself impaired, while inconsistent regional and national messages congested the alert chain. [2][5][7][18]
  • KPN reported restoration work and corrective measures, including software-configuration changes and a faster alternative path for 112 traffic if the routing platform slowed or stopped. Regulators also called for continuous testing across the complete 112 chain, stronger change resilience and attention to identical software and shared dependencies. Statements that measures were accepted or implemented are not, by themselves, public proof of long-term effectiveness. [2][7][9][10]
  • Responsibility follows practical control. KPN controlled routing-platform architecture, software change, monitoring and restoration evidence. Other operators controlled their dependency awareness and customer-facing continuity. Government, police, safety regions and care organisations controlled fallback plans and public instructions. Regulators controlled requirements, investigation and follow-up. A credible closeout must show how these controls work together when the normal national call path fails.

A national service depended on one operator's call path

The starting point for assigning responsibility is not the software defect. It is the route that an emergency call had to take before the defect mattered.

The joint investigation describes a centralized chain. A person called 112 from a fixed or mobile voice service. Traffic from all voice providers was sent through KPN's telephone network to the public safety answering point in Driebergen. An operator there answered and transferred the call to the appropriate regional emergency response centre, which then alerted the required service. The head of the police was controller for the 112 domain, while the Minister of Justice and Security was responsible for the chain. [2][7]

That architecture distributed institutional responsibility but concentrated a critical transport function. A non-KPN subscriber could have a commercial relationship with another operator and still depend on KPN for the final national route into 112. The public-safety consequence of a KPN routing failure was therefore not bounded by KPN's retail customer base. This is why the incident should not be analysed as an ordinary supplier outage measured only by the percentage of one company's subscribers who could place calls.

The dependency also changes the meaning of redundancy. A mobile provider may have diverse radio sites, transport links and core components within its own network. A fixed operator may have separate access technologies. Those design choices do not establish emergency-call resilience if all paths converge on the same downstream routing platform. Independence has to survive the whole service path. Diversity before a convergence point can improve ordinary availability while leaving a common national failure domain intact.

There is nothing inherently irresponsible about a shared national ingress. Centralization can simplify call handling, location, transfer and operational coordination. It can also make specialist controls easier to maintain. But concentration raises the required standard of evidence. The operator of the shared point must show that redundancy is not only physical, that state does not silently synchronize across nominally separate systems, that failure cannot be amplified by repeated demand, and that a route around the platform remains available under realistic load.

The public authorities responsible for the chain face a parallel duty. They need an architecture map that identifies where contractual diversity ends and technical convergence begins. They need to know which organisation can reroute traffic, which organisation can declare that a fallback is safe, and which tests exercise calls from every originating network through to a human answer and regional transfer. Without that end-to-end view, each entity can report that its own component is available while the public service remains unreachable.

The 2019 outage made this control gap visible. KPN owned and operated the platform whose failure stopped ordinary forwarding. Yet the societal dependency was larger than KPN, and the ability to mitigate it was divided among operators, police, ministries, safety regions, emergency services and care organisations. Accountability therefore cannot be reduced to either "KPN's bug" or "government preparedness." The architecture created distinct control obligations for both.

The timeline separates detection, diagnosis, restoration and public recovery

The supervisory record supplies a more useful chronology than a single outage-duration figure.

At 15:32, KPN's monitoring centre received the first report of a decline in visible traffic. More reports followed. The joint report places the principal malfunction from 15:34. KPN used monitoring signals, customer responses and reports from its own organisation to initiate an emergency procedure. The first signal and the defined start of broad service loss were therefore close in time, but detection of abnormal traffic was not the same as knowing the mechanism or restoring a route. [2][7]

At 17:45, the investigative team identified the cause. At 18:30, KPN successfully restarted the first system. At 18:52, voice service and access to the public safety answering point were restored. Those timestamps distinguish several operational questions. Monitoring saw a symptom quickly. Technical diagnosis took roughly two hours from the first visible decline. A restart began recovery, but full accessibility came later. Each interval belongs to a different control surface: telemetry, incident escalation, fault isolation, safe restart and service verification.

Public recovery extended beyond network restoration. Government bodies, police, safety regions, emergency services and care organisations scaled up crisis operations and tried to offer alternatives. By 21:00, organisations had scaled down their crisis structures. A final NL-Alert message reporting resolution was sent at 21:30. The network's return at 18:52 did not instantly remove the need to reconcile public instructions, reopen normal contact paths and withdraw temporary arrangements. [2][7]

This chronology matters because an availability metric can hide the shape of the response. A three-hour outage may appear to sit within a four-hour internal standard, but an emergency service is not adequately measured by duration alone. The number of affected people, the national scope, the criticality of calls, the loss of alternative service numbers and the confusion around fallback instructions all change impact. KPN's own 2019 reporting later questioned whether its weighted-downtime performance indicator sufficiently represented incidents with severe societal consequences. [10]

The timeline also exposes evidence that remains unavailable publicly. The report does not provide every alarm, counter value, operator command, escalation decision or restart criterion. It does not identify an individual engineer or software supplier as the decision owner. It does not show the complete event log from 15:32 to 17:45. Those gaps do not erase the established mechanism, but they limit claims about why diagnosis took as long as it did and whether a different operational choice would have shortened the incident.

A rigorous closeout would preserve this separation. Detection evidence would show when the first meaningful threshold crossed and whether staff understood that 112 was affected. Diagnosis evidence would show how responders distinguished load from state corruption. Restoration evidence would show why restarting one system was safe and how traffic was controlled. Service evidence would show successful calls from each operator, not only that platform processes were running. Public recovery evidence would show when accurate, consistent alternatives reached citizens. "Resolved" should be the last of those tests, not the first.

The triggering defect was only one layer of the root cause

The joint report gives unusual specificity about the routing-platform failure while also showing why "software bug" would be an incomplete explanation.

KPN's call routing platform was an essential part of the telephone network. It supplied the information needed to place each call on the correct route, including calls to 112. The platform contained four call routing systems described as independently operating. The direct failure involved software configuration, synchronized operation and counters used to monitor routing requests. [2][7]

The report describes a chain. A software change in the platform's service-management system unintentionally caused the counters of the four routing systems to run synchronously. A separate script implemented in January 2019 was intended to warn when the counters reached 95 percent of their maximum, but an implementation error meant the counters were not reset in time. On 24 June, all four counters reached a negative value at nearly the same moment. That state generated a large volume of error messages. Every new routing request generated another error, and repeated call attempts increased traffic.

After about an hour of accumulating error and request load, the platform could no longer process call-routing requests. [2][7]

This chain contains at least four analytically different elements.

The triggering condition was the counters crossing into the negative state. The latent technical defect was the software and configuration behaviour that allowed the counters to synchronize and fail together. A detection control failed because the warning-and-reset script did not operate as intended. An amplification condition arose because each routing request produced another stored error while callers naturally retried. The architectural consequence was that four systems presented as independent no longer provided useful failure isolation.

The report adds another dependency decision. In June 2018, KPN began using the call routing platform to route 112 traffic while the 112 platform was being upgraded. When the routing platform failed, information needed to route emergency calls was unavailable. That decision did not create the counter defect, but it connected the platform's failure to national emergency access. It is therefore a contributing architectural condition rather than the immediate trigger. [2][7]

Separating these layers prevents accountability from collapsing onto the last visible error. A counter can overflow or become negative because software behaves incorrectly. But an organisation decides how state is partitioned, what alerts are tested, whether identical systems share a management plane, what happens when error storage grows under demand, and whether emergency traffic has a route that does not depend on the same logic. A human mistake in a script may be real without being a complete root cause.

The same discipline protects against unsupported claims. The public record does not name the vendor, the exact code module, the counter width, the maximum value, the engineer who wrote the script or the approval process for the service-management change. It would be tempting to make the mechanism feel more technical by supplying those details, but doing so would replace evidence with invention. The established conclusion is narrower and still significant: supposedly independent routing systems shared state behaviour that defeated redundancy, and an intended warning control failed to prevent synchronized exhaustion.

Four systems were not four independent failure domains

Redundancy is often communicated as a count. Four systems sound safer than one. The KPN incident shows why the number of components is not the same as the number of independent failure domains.

The call routing systems could operate separately in a normal sense and still share the properties that mattered during this event. They used identical or closely related software, were affected by a common service-management change, maintained counters that became aligned and reacted similarly when those counters crossed the failure condition. Their physical or process separation did not stop a common state transition. Once the same failure reached all four, the architecture's redundant capacity could not carry traffic around the problem. [2][7]

This is common-mode failure: multiple components fail because they share a cause, dependency, state or assumption. The term should not be used as a vague synonym for "large outage." It identifies why redundancy did not reduce the probability or impact as expected. In this case, synchronization changed the risk. Counters that would have reached a problematic state at different times might have produced a warning, a partial failure or an opportunity to reset one system while others continued. Alignment converted staggered exposure into near-simultaneous loss.

The error-storage mechanism made the common mode operationally worse. As calls were retried, more requests generated more error messages. A service under public stress experienced demand that was both legitimate and predictable: people call again when a call does not connect. A design that turns retries into accumulating internal error work can move away from recovery as users seek help. Load shedding, bounded logging, backpressure and emergency-traffic prioritisation are therefore not generic performance features here. They are part of safety continuity.

The regulator's recommendations explicitly broadened the lesson beyond KPN. The telecom sector was told to identify new weaknesses and dependencies involving operational systems, database connections, configuration changes, software updates and identical software. That list is an architecture test. It asks whether supposedly redundant elements share the same data store, control plane, release package, operational procedure or failure-sensitive state. [2][7]

Evidence of genuine independence would be concrete. It could include staggered state, separate management domains, version diversity where justified, bounded failure effects, an emergency route that bypasses the impaired platform, and tests that inject the exact common-mode conditions. It would also include organisational independence: authority to isolate a system, reroute traffic and stop a change without waiting for the same team or tool that is failing.

None of those controls should be assumed from a diagram showing four boxes. Nor should they be inferred from the statement that systems are redundant. The evidence should demonstrate what happens when a common management change is wrong, when counters reach boundary conditions together, when error logging accelerates and when callers retry at national scale. The burden is especially high when the platform carries emergency traffic from other operators.

A fallback is independent only if it avoids the failed assumptions

KPN reported that it improved resilience by enabling 112 traffic to be rerouted quickly through alternative channels if the routing platform became slow or stopped. That is directionally responsive to the incident. It addresses the need to route around the platform rather than simply restart identical systems. The statement still leaves important questions about independence and proof. [2][7]

A fallback that uses the same control plane, service-management software, database, routing data or operational approval path may be alternative in topology but common in failure. If the primary and backup read the same corrupted state, depend on the same counters or require commands from an impaired management system, switching paths does not remove the cause. The incident makes "alternative" a claim that has to be decomposed.

Technical independence asks whether the backup can determine and forward the correct emergency route without the failed platform. Capacity independence asks whether it can carry national retry demand, not merely a small test call. State independence asks whether it maintains or receives routing information through a separate mechanism. Control independence asks whether operators can invoke it when ordinary management tools are degraded. Organisational independence asks who has authority to switch and whether that authority is available around the clock.

Time also matters. A fallback that exists but requires two hours of diagnosis before activation may reduce recovery time only after responders understand the cause. A safer design can use observable service criteria: if end-to-end emergency-call success falls below a threshold, route traffic away from the platform even before the exact defect is known. That approach creates its own risks, including false switching and overload, so it must be tested. But it moves continuity from fault diagnosis toward service outcome.

The June 2018 decision to route 112 through the call routing platform is relevant here. A temporary or migration-related dependency can become a durable production assumption. Upgrade programmes should therefore carry an explicit expiry and verification record: why the dependency was introduced, when it will be removed, what failure modes it adds, and what route remains if the interim component fails. The public report does not disclose the complete decision record, so it cannot establish whether those controls existed. It does establish that the platform's failure made 112 unreachable.

Independent fallback also extends beyond KPN. Other operators need to know whether they can deliver emergency calls without the shared route and under what conditions. Public authorities need alternatives that do not assume ordinary voice service. Care organisations need communications tools whose users are trained and whose dependencies are understood. A backup network that staff do not know how to use is not operationally independent, even if its technical path is separate.

The appropriate post-incident question is therefore not "Was a backup added?" It is "Which failed assumptions does the backup avoid, and what evidence shows it can take national emergency traffic while the primary platform, management plane and normal communications are impaired?"

End-to-end testing was a governance control, not a final technical check

The joint report identified a lack of end-to-end service management in the 112 chain and recommended continuous testing and monitoring across the entire path. It noted that KPN continuously tested 112 routing in the TDM network with a call generator, while no equivalent method was available for the mobile network after the upgraded 112 platform was implemented. [2][7]

That finding is central to accountability. Component tests could show that an originating network accepted a 112 call, that a KPN router was healthy, that the answering point could receive a test input, or that a regional centre could take a transfer. None proves that a real call from each provider traverses all dependencies and reaches the intended human endpoint. The public service is the chain, not any individual component.

Continuous testing does not necessarily mean placing audible test calls into emergency operations without controls. It means creating a safe method that exercises signalling, routing, transfer and observability without confusing operators or the public. Synthetic transactions can be marked, rate limited and directed to controlled endpoints. The technical design is important, but so is ownership. Someone must decide which origins are tested, who receives failures, how quickly an alert is escalated and when a failed test triggers a continuity action.

Coverage should follow the architecture. Tests need to originate from each mobile and fixed provider, from relevant access technologies, and from conditions that expose convergence. They should verify ordinary routing and alternative paths. They should exercise changes before and after deployment, as well as long-running state that cannot be reproduced by a brief functional check. Boundary testing should include counter exhaustion, synchronized state, error-volume growth and the effect of retries.

The result should be measured as an emergency-service outcome. Did the call reach the national answering point? Was caller information handled as expected? Could the call be transferred to the correct region? Was the round-trip time acceptable? Did monitoring associate a failure with the correct dependency? A platform health dashboard that remains green while end-to-end calls fail is not meaningful assurance.

Governance enters because the chain crosses organisations. KPN could test what it controlled, but the Minister, police, other operators, safety regions and emergency services controlled other parts of the path. No single component owner could certify the whole service without cooperation. The regulator's recommendation therefore implied a shared operating model: agreed test cases, common thresholds, evidence retention, escalation duties and authority to demand remediation.

Publication of all sensitive test detail would be inappropriate. But aggregate evidence could be public without revealing exploitable architecture: coverage by operator and access type, test frequency, failure rates, maximum detection time, fallback exercise dates and closure of material findings. Such evidence would let regulators and the public distinguish an accepted recommendation from a working assurance programme.

Monitoring saw traffic decline but missed the condition that mattered

KPN's monitoring centre received a signal at 15:32, close to the reported start of the broad malfunction. That is evidence that some observability worked. The harder question is whether the organisation monitored the leading condition and the public service outcome.

The counters had approached a maximum before becoming negative. A script was intended to warn at 95 percent and support a timely reset, but an implementation error prevented the control from doing its job. This was not simply a failure to notice that customers could not call. It was a failure of a specific preventive signal that should have surfaced dangerous state before the platform stopped processing requests. [2][7]

The distinction matters for incident economics. Detecting a nationwide traffic decline after failure has begun can shorten restoration. Detecting synchronized counter growth before boundary crossing can prevent the incident. Monitoring budgets and operational attention should therefore be judged by the control they enable. A dashboard metric has limited value if it cannot cause an operator or automated system to act safely before impact.

The report also found limited public evidence exchange of specific performance indicators between network elements that prevent overload. That suggests another boundary: local components may have known about queue, error or capacity pressure without turning it into an end-to-end service signal. A complex network needs both local diagnostics and service-level synthesis. Local detail supports diagnosis; service outcome supports prioritisation.

Retry traffic was predictable. When a call fails silently or does not connect, callers try again, and institutions may originate additional calls while checking service. Monitoring should distinguish original demand from retry amplification and should anticipate that public concern will increase load. A system whose error path stores work for every retry needs particularly strict bounds and alerts.

An accountable monitoring closeout would show at least four layers. Preventive telemetry would show counters, state alignment and boundary conditions. Platform telemetry would show routing success, error rates, queue and storage pressure. Service telemetry would show successful end-to-end 112 calls from each operator. Societal telemetry would show whether fallback numbers and public instructions were being used successfully. These layers support different decisions and owners.

The public record does not show the precise thresholds introduced after the event or the complete alarm history. It should not be assumed that one failed script represented all monitoring. But the established gap is sufficient to reject a simple claim that fast initial symptom detection proves adequate control. The incident began close to the first reported traffic decline because an earlier preventive control had not constrained the synchronized state.

The separate NL-Alert fault tested the independence of public warning

NL-Alert failed for a different reason. On 24 June, a configuration change connected to 4G reporting and a periodic network scan overloaded an adapter in KPN's Cell Broadcast platform. KPN could not process NL-Alert messages over 4G until the problem was identified and resolved the next day. The joint report explicitly treated this as separate from the telephone routing failure. [2][5][7]

That causal boundary must be preserved. The telephony counters did not cause the Cell Broadcast adapter problem. The fact that both involved software or configuration does not make them one incident mechanism. Combining them would distort technical accountability and could assign corrective actions to the wrong control.

The simultaneous effect is nevertheless relevant to infrastructure resilience. Public authorities used NL-Alert as one way to tell people that 112 and the national police service number were unavailable and to provide alternatives. KPN customers on 4G did not receive those messages through the expected path. Other problems then affected the broader alert process: regional and national messages were numerous and inconsistent, the central chain became congested, some messages arrived very late, and one national message included an incorrect number associated with a newspaper tip line. [2][7][18]

This was a continuity-of-continuity problem. A warning system used when ordinary communications fail must have failure assumptions that differ from the service it supports. Cell Broadcast is technically distinct from a voice-routing platform, but both still depended on operator infrastructure, configuration practice, monitoring and coordinated public content. Technical diversity alone did not guarantee usable warning.

There are at least three independence tests. The delivery path must survive the incident it is supposed to explain. The control path used to create and send messages must remain available and understood. The information process must produce one clear, verified instruction rather than competing alternatives. Failure in any one can make the warning ineffective even if the others work.

The report found that KPN did not detect the 4G NL-Alert problem quickly enough and that NL-Alert was not treated as a separate critical service inside KPN. KPN later added monitoring and included network-scan behaviour in testing. Those measures address the technical path. The public authorities also needed procedures for a national 112 outage, consistent message ownership and usable alternatives. [2][7]

This division prevents misplaced blame. KPN could not decide every regional instruction, and safety regions could not repair the 4G adapter. KPN controlled platform detection and delivery. Government actors controlled message governance. Both had to work for the public warning function to succeed.

Crisis plans existed, but many were not operational

The Netherlands did not enter the incident with no continuity policy. Agreements had followed earlier 112 disruptions, and the police maintained Generic Operational Scenarios. Scenario 4 came closest to a loss of the public 112 infrastructure and included staffing police and fire stations. A 2013 government letter also offered actions for citizens, such as trying a mobile phone if a fixed call failed, trying a fixed phone if mobile failed, or going to an emergency-service location if telephone facilities were unavailable. [2][7][18]

The investigation found a gap between documentation and operational readiness. Security regions had not been fully involved in the earlier action framework. Roles, communication methods and implementation details were incomplete. Some organisations had little knowledge of the documents. Plans often assumed a regional incident, not national unavailability. Scenario 4 also assumed that the national police hotline 0900-8844 would work, but the same KPN failure made that number unavailable. [2][5][7]

That is an infrastructure lesson. A fallback instruction must be checked against the same dependency map as the primary service. Offering another telephone number is not meaningful if it enters the same failed routing platform. Advising people to visit a station can work only if the public knows which locations are staffed and those locations have functioning communications with dispatch. A plan can be formally approved while its operational prerequisites remain unspecified.

During the event, organisations improvised. Police and fire stations were made available in some places, extra staff were deployed, social media was used, and local alternatives were announced. Resourcefulness reduced some consequences, but improvisation also produced inconsistency. The Ministry delayed a uniform national message while seeking broader options because 0900-8844 was unavailable. The joint report concluded that this delay contributed to loss of control over crisis communication. [2][7]

The lesson is not that every crisis can be scripted. It is that the stable parts should be pre-resolved. Message authority, verification of alternative numbers, location data for staffed stations, communications among national and regional actors, and criteria for using NL-Alert can be agreed before an outage. Exercises can reveal whether staff know the plan and whether the fallback shares the failed network.

Plans should also state their assumptions. If an action relies on mobile data remaining available, that should be explicit, along with an option for incidents in which it is not. In 2019, KPN Internet services continued to work, allowing web-based communications for some users. That fact made tools such as WhatsApp or Skype useful in parts of the response, but it should not be generalized into a universal emergency substitute. It assumes data access, a compatible device, a reachable destination and users who know what to do.

Operational readiness is therefore measured in observed capability, not document count. Can staff initiate the procedure? Do alternatives avoid the primary failure? Can the public understand one verified instruction? Can care organisations contact partners? Are fallback systems exercised often enough that personnel remain familiar? The report's recommendations focused on implementation, familiarity and compliance because the policy layer alone had not produced those outcomes.

Public-safety harm must be measured without inventing causation

The supported harm was serious. People could not use the ordinary national emergency number, the police service number was also unavailable, alternatives differed by region, alert delivery was impaired and care organisations had to improvise communications. The report describes gaps in emergency healthcare and substantial societal impact. Those findings justify a high-impact assessment without needing a dramatic but unproved casualty claim. [1][2][5][7]

The joint investigation discusses three deaths reported by regional ambulance services during the outage period. It also says the services responded within applicable time limits and protocols, and the health inspectorate could not establish whether delayed initiation of paramedic assistance played a role in the deaths. A separate hospital transfer was delayed by 20 minutes, but the hospital's review found no direct consequence for that patient. Another complaint led to improvement points without established patient harm. [2][7]

These distinctions are essential. "People died during the outage" is a temporal statement. "The outage caused deaths" is a causal statement that the cited investigation did not establish. Repeating the first in a context that implies the second would overstate the evidence and could distort both public understanding and legal exposure.

Uncertainty does not make the incident harmless. Emergency communications are designed for situations in which delay can matter even when a later investigation cannot reconstruct a counterfactual outcome. The accountable measure is exposure: how many call attempts failed, how long callers waited, which alternatives were available, whether care organisations lost contact paths and whether responses began later than they otherwise would have. The public report provides examples and institutional findings but not a complete call-attempt dataset.

This points to an evidence requirement for future incidents. Operators and public authorities should preserve privacy-protective records that can connect failed call attempts, retry patterns, alternative contacts and dispatch timing. Such analysis must be carefully governed because emergency-call data is sensitive. Aggregate figures and controlled investigations can still establish whether failure concentrated by region, provider, access technology or time.

Impact measures should also distinguish reachability from responsiveness. Restoring the ability to dial 112 does not prove that every queued or retried demand was handled normally. Conversely, an unanswered call may have reasons outside the routing incident. The goal is not to assign every outcome to the network but to quantify the additional risk created by loss of the ordinary path.

KPN's stated concern about weighted downtime is relevant because conventional availability measures can discount exactly this kind of event. A short but nationwide failure of a critical service may create greater public risk than a longer partial failure of a less consequential feature. A useful metric should therefore include service criticality, population affected, cross-operator reach, fallback availability and the time required to restore end-to-end success. [10]

Those measures should inform investment and accountability before an incident, not only its retrospective description. If emergency-call continuity receives a higher risk weight, common-mode testing, independent fallback and continuous chain monitoring compete more effectively for engineering resources. The metric then becomes a governance tool rather than a public-relations number.

KPN's corrective actions addressed the mechanism, but proof requires more than acceptance

The joint report says KPN carried out a root-cause analysis and extensive evaluation and commissioned Bell Labs Consultancy. KPN formulated an action plan in August 2019. The report states that most measures had been implemented by the time of publication. It identifies software-configuration adjustments intended to prevent routing requests from being disrupted by large volumes of error messages and a quicker alternative channel for 112 traffic when the routing platform slowed or stopped. [2][7]

These measures map sensibly to the failure. Preventing error-message accumulation addresses amplification. Changing counter and configuration behaviour addresses the synchronized state. Alternative routing addresses the loss of the platform. Additional monitoring addresses detection. Treating 112 as a distinct critical service gives the chain clearer internal priority.

The regulator concluded that the action plan would make the network more robust and reduce recurrence risk. It also found that limited public evidence attention had been paid to planned and unplanned software-configuration vulnerabilities, change resilience, performance-indicator exchange, end-to-end service management and process discipline. It recommended periodic progress reporting and said regular supervision would examine compliance. [2][7][9]

Those are meaningful supervisory findings, but they are not the same as public evidence that every control remained effective over time. "Implemented" can mean that a configuration was changed or a procedure adopted. "Effective" requires a test showing the control prevents, detects or limits the relevant failure. "Sustained" requires evidence after further software releases, platform migrations and personnel changes.

A strong remediation record would bind each action to a test. The counter correction would be tested at boundary values and under synchronized state. Error handling would be subjected to repeat traffic and bounded-storage conditions. Alternative routing would be exercised while the primary platform and its management dependencies were unavailable. Continuous end-to-end tests would cover every originating operator. Monitoring would demonstrate detection of both platform degradation and actual call failure. Crisis exercises would test a single national message and verified non-voice alternatives.

The results should include failure, not only success. A test programme that never finds a defect may have weak coverage. Useful evidence records what was injected, what signal appeared, which action followed, whether traffic remained available and what was repaired before the next exercise. It also records limitations: a laboratory load may not represent national retries, and a synthetic call may not exercise every handoff used in production.

KPN's annual report provides the operator's own account of restoration, stabilisation and improvement. It is relevant because it shows what management chose to disclose and how the company framed the effect. It should not be treated as independent verification. The joint supervisory report and later regulator follow-up provide a separate layer, but even they do not publish every test result or internal change record. [9][10][11][12]

The appropriate conclusion is therefore calibrated. Public evidence supports that KPN took substantial remedial action and that regulators reviewed and monitored it. Public evidence does not support saying recurrence became impossible, every fallback was independently verified at national load or all long-term residual risk was eliminated.

Compliance was a floor, not a proof that the architecture was adequate

The Dutch Telecommunications Act and associated continuity rules required providers of public electronic communications networks and public telephone services to take appropriate technical and organisational measures, maximise availability during technical or power failures and report significant continuity interruptions. Dutch policy also addressed the ability to reach 112 through mobile services. At EU level, Article 109 of the European Electronic Communications Code required access to emergency services through the single European number 112 without charge. [13][14][15][16]

The joint investigation found that KPN complied with the continuity obligations it examined, while also finding that the outage occurred despite compliance. That combination is important. It prevents two simplistic conclusions.

First, the incident is not evidence by itself that KPN violated every applicable continuity rule. A regulator assessed the legal obligations and did not make that finding in the cited report. A responsible article should not convert an outage into a legal verdict.

Second, compliance did not demonstrate that the system could withstand the actual common-mode condition. General duties such as appropriate measures and maximum availability require judgement. They cannot enumerate every interaction among identical software, synchronized counters, error storage, repeat calls and a national emergency dependency. A company can satisfy the assessed baseline and still discover that its architecture contains a material untested failure mode.

This is why regulatory accountability should include evidence quality. Requirements should ask not only whether a continuity policy exists but how the operator established independence, which end-to-end tests ran, how changes affected emergency routing and what residual risk remained. The answer may still be risk based rather than absolute. No network can promise zero failure. But the decision to accept residual risk should be visible to the responsible authority and supported by tests that reflect the service's public importance.

Incident notification is another control. Timely notification enables regulators and government actors to coordinate response and preserve evidence. It does not substitute for public instructions. A provider can notify an authority while citizens still receive inconsistent alternatives. Legal reporting, crisis communication and technical restoration are related but separate duties with different audiences.

The event also illustrates why regulation must follow service chains rather than corporate boundaries. Other operators originated calls, KPN transported them into the 112 path, police controlled the answering domain, the ministry held chain responsibility, safety regions acted locally and care organisations depended on communications. A requirement applied to only one entity cannot create end-to-end assurance unless the interfaces and shared tests are also governed.

Regulators can make that assurance more auditable by requesting stable indicators: successful emergency-call tests by origin, maximum detection time, time to invoke fallback, unresolved high-risk change findings and dates of national continuity exercises. Sensitive details can remain protected while trends and material exceptions are disclosed.

Compliance is therefore necessary but not dispositive. It sets a minimum expectation and a mechanism for intervention. The 2019 report shows that accountability still requires examining whether the implemented controls matched the real architecture and whether the evidence could detect a failure that the rulebook did not name in advance.

Emergency-session standards provide context, not proof of KPN's exact design

ETSI and 3GPP specifications describe emergency sessions in IP Multimedia Subsystem environments, including functions used to recognise, route and handle emergency communications. They are useful context because modern voice networks increasingly implement service logic in software and depend on control functions that may be virtualised, replicated and managed centrally. [17]

The standard should not be used to claim that KPN's 2019 routing platform had a particular IMS component, interface or deployment topology. The source set does not establish that mapping. A standards diagram is not an incident architecture diagram.

The useful lesson is methodological. Emergency communication is a service outcome assembled from multiple functions: identifying an emergency request, selecting a route, transporting it, reaching the appropriate answering point and supporting transfer. Redundancy at one function does not guarantee the outcome if another function is common. Software replication can increase availability while also reproducing the same defect and state.

Virtualisation makes this issue more important, which is why the regulator recommended controls for software and configuration errors in anticipation of increasing network virtualisation. A virtual network function can be created quickly and moved between hosts, but copies may share the same image, orchestration, policy, database and management credentials. Physical dispersion can coexist with logical common mode. [2][7]

Standards compliance likewise cannot replace service testing. A component may implement its specified interface correctly while the production chain fails because routing data is unavailable, management state is corrupted or another operator's traffic is not covered by the test. Interoperability testing establishes one kind of assurance. Continuous end-to-end monitoring establishes another.

The standard context therefore sharpens the questions without answering them. Which functions were in KPN's call path? Which were common across the four routing systems? Which state was shared? Which alternative path bypassed those functions after remediation? The public record answers the broad routing-platform mechanism but not a complete implementation inventory.

This boundary protects technical accuracy. It would be easy to use standards terminology to make the account sound precise. Unless a source binds that terminology to the incident, it can create false confidence. The correct use is to explain why emergency service depends on a chain of functions and why replicated software requires common-mode controls, while leaving KPN's exact unpublished topology unresolved.

Accountability follows control over prevention, detection, containment and proof

Responsibility for the outage was distributed, but it was not vague. Each actor controlled identifiable parts of prevention, detection, containment, communication, restoration and verification.

Control area Primary practical controller Evidence expected
Routing-platform architecture KPN Dependency map, failure-domain analysis, common-mode test results and change records
Counter and long-running-state safety KPN and relevant supplier Boundary tests, warning-control verification, reset logic and ownership of corrective action
Error amplification and overload KPN Bounded logging, retry-load tests, backpressure behaviour and service-level alarms
112 alternative routing KPN with chain authorities Proof the fallback bypasses failed dependencies, capacity tests and activation records
Cross-operator delivery KPN, other operators and chain authorities End-to-end call tests from every originating network and access class
National 112 governance Minister of Justice and Security and police controller Current architecture ownership, decision rights, exercise records and escalation criteria
Regional fallback operations Police and 25 safety regions Staffed-location procedures, verified alternatives, training and exercise outcomes
Care continuity Ambulance, GP, hospital and regional health organisations Scenario playbooks, independent communication capability and staff familiarity
NL-Alert technical delivery KPN and other mobile operators Continuous non-disruptive monitoring, configuration tests and delivery evidence
Crisis-message governance Ministry, police and safety regions Single message authority, verified numbers, timing records and correction procedure
Legal oversight and follow-up Dutch supervisory authorities Progress reports, inspection findings, residual-risk decisions and closure evidence

This map prevents two opposite errors. One is to blame KPN for every confused public message, even though government and regional bodies controlled message content and execution. The other is to diffuse the routing failure across the whole chain until no actor remains responsible for the platform. KPN had practical control over the call routing system, its change process, monitoring and technical fallback. That responsibility remains specific even when other actors also had continuity duties.

Control also determines what evidence can reasonably be demanded. Citizens cannot produce platform counter logs. Other operators cannot independently prove how KPN's four systems managed state. KPN cannot prove that every safety region trained its staff. Each controller should supply the records within its authority, while the chain owner assembles them into an end-to-end case.

Suppliers may share technical responsibility, but the available sources do not identify the supplier responsible for the relevant software or configuration. It would be improper to assign a vendor fault without evidence. Contracting does not remove KPN's operational responsibility to test and monitor a critical platform, just as operator control does not automatically prove that KPN authored every defective component.

The regulator's role is not simply to declare recommendations accepted. It can test whether risk controls are measurable, whether progress reports bind to current systems and whether major changes reopen closed findings. If the routing platform is replaced, remediation evidence tied only to the old platform may no longer assure the service. Oversight should follow the continuing emergency function.

This approach also makes accountability constructive. It does not require identifying a person to punish before controls can improve. It asks who could change the condition, who could see it, who could limit impact and who can verify the repair. Where those answers are missing, the absence is itself a governance finding.

A credible closeout would show independence over time

The public record establishes a mechanism and a set of responses. The remaining question is what would justify closing the risk.

First, KPN would need current architecture evidence. That includes the ordinary 112 route, the alternative route, management dependencies, routing-data sources and convergence points used by traffic from other operators. The purpose is not to publish a sensitive network blueprint. It is to allow authorised reviewers to test whether the fallback avoids the platform and state that failed.

Second, the operator would need change evidence. The service-management update that aligned counters and the warning-script error show why functional release testing was limited public evidence. Reviews should cover long-running state, boundary values, synchronization across replicas and the behaviour of old state after an upgrade. They should also establish who can stop a release when emergency continuity evidence is incomplete.

Third, the chain would need repeated end-to-end tests. One successful test after remediation would show that the route worked once. It would not show that every operator, access technology and fallback remained covered after later changes. Continuous or frequent testing, with controlled synthetic calls, can detect regression. Periodic national exercises can test the organisational layer that synthetic calls cannot.

Fourth, fallback evidence would need realistic failure injection. The primary platform should be made unavailable in a controlled environment or exercise. Management services, routing data and normal communications should also be constrained where safe. The alternative path should carry representative load, and responders should activate it using the same authorities and tools available during an incident.

Fifth, public communication should be exercised as infrastructure. Message templates need verified alternatives that do not share the failed route. National and regional actors need a process that prevents conflicting numbers and alert congestion. Staff should know when one national instruction takes precedence and how corrections propagate.

Sixth, impact metrics should reflect societal service. Availability, successful call completion, detection time, fallback activation time, population affected and cross-operator scope belong in the performance picture. KPN's concern about weighted downtime was a useful acknowledgement that ordinary network metrics may not capture critical-service impact. [10]

Seventh, independent follow-up should record residual risk. Some common-mode conditions may be reduced rather than eliminated. A reviewer should state which dependencies remain, why they are accepted, what detects them and when the decision will be revisited. Silence should not be interpreted as zero risk.

The later regulator report says KPN accepted recommendations and that follow-up was monitored. That supports an account of continued oversight. The available public packet does not include every periodic progress report or current test result. The right conclusion is therefore not that remediation failed, but that public proof is incomplete. [9]

This standard may appear demanding for an event from 2019. The service, however, is enduring. Emergency networks evolve through virtualisation, supplier changes, platform upgrades and new access technologies. Evidence that was persuasive immediately after one incident can become stale. Closeout must be a maintained assurance process rather than a one-time declaration.

What new evidence could change this assessment

Several conclusions could become stronger or narrower if additional records were made available.

Complete platform logs could establish the exact sequence from counter alignment to error accumulation and show whether alarms fired before visible traffic declined. Software-change and approval records could identify which tests were required and which teams controlled the risk. A supplier root-cause analysis could clarify component ownership without speculation.

Pre-incident and post-remediation failover tests could show whether an alternative route existed before 24 June and how its independence changed afterward. End-to-end records could establish coverage across operators, fixed and mobile access, the answering point and regional transfer. Capacity exercises could show whether the fallback can handle retry demand.

Call-attempt and completion data could improve impact measurement. Properly protected, it could show how many calls failed, how retry behaviour evolved and whether service returned uniformly. Care-sector records could clarify operational delays while preserving the report's caution about individual outcomes.

Periodic regulator findings could show whether KPN completed the action plan, whether controls remained effective after subsequent changes and what residual risks were accepted. Aggregate public indicators could provide assurance without exposing sensitive details.

Evidence could also narrow responsibility. If a supplier contract and technical record showed that a component behaved contrary to specification despite reasonable testing, supplier accountability would become more specific. If internal records showed that a known warning failure was accepted without mitigation, management responsibility would become more specific. The current source set supports neither claim.

The assessment should therefore remain provisional at the edges and firm at the centre. The outage window, national dependency, broad voice impact, continued Internet availability, four-system common mode, failed counter warning, separate NL-Alert mechanism and preparedness gaps are well supported. Individual causation, vendor identity, internal decision ownership and complete long-term effectiveness remain unresolved.

Conclusion: network resilience has to survive the shared path

The KPN outage became a public-safety accountability test because the Netherlands' ordinary emergency-call path converged on one operator's routing platform. Four routing systems did not provide four useful failure domains once software state synchronized. A preventive warning did not stop counters crossing the boundary. Repeated call demand amplified error work. The routing platform stopped forwarding calls, and the failure reached 112 traffic from other operators.

The incident also showed that technical restoration is only one part of continuity. A separate NL-Alert fault impaired one warning channel. Government and regional plans were not consistently operational. Alternative numbers and instructions varied. Care organisations relied on improvisation and communications tools that were not always familiar. These were distinct failures with distinct controllers, but they combined in the public experience.

KPN's remedial actions addressed important parts of the mechanism, and regulators established follow-up. That evidence supports neither dismissal nor certainty. It supports a verification agenda: prove that routing fallback avoids the failed assumptions, continuously test the full multi-operator 112 chain, monitor service outcome as well as platform state, bound retry amplification, rehearse public alternatives and maintain evidence after every material change.

Accountability is clearest when it follows control. KPN controlled the platform and its technical repair. Other operators controlled their awareness of and testing against the shared path. Police and the ministry controlled the national chain. Safety regions and care organisations controlled local continuity. Regulators controlled the standard of proof and follow-up.

The lasting lesson is not that redundancy failed despite four systems. It is that redundancy had been counted at the component level while risk accumulated at the shared-state and service-chain levels. For emergency network infrastructure, independence is not a label on an architecture diagram. It is an outcome demonstrated under the exact conditions that could make every normal path fail together.

Sources

  1. https://www.rdi.nl/documenten/2020/06/25/onbereikbaarheid-van-112-op-24-juni-2019
  2. https://www.rdi.nl/site/binaries/site-content/collections/documenten/2020/06/25/onbereikbaarheid-van-112-op-24-juni-2019/Gezamenlijk%2Brapport%2B112%2BAT%2BIJenV%2Ben%2BIGJ%2Bonbereikbaarheid%2Bvan%2B112%2Bop%2B24%2Bjuni%2B2019.pdf
  3. https://www.inspectie-jenv.nl/actueel/nieuws/2019/06/26/onderzoek-naar-storing-112
  4. https://www.inspectie-jenv.nl/actueel/nieuws/2019/08/22/plan-van-aanpak-onderzoek-112-gepubliceerd
  5. https://www.inspectie-jenv.nl/actueel/nieuws/2020/06/25/overheden-en-organisaties-niet-voldoende-voorbereid-op-landelijke-uitval-112
  6. https://www.inspectie-jenv.nl/documenten/2020/06/25/rapport-onbereikbaarheid-van-112-op-24-juni-2019
  7. https://www.inspectie-jenv.nl/site/binaries/site-content/collections/documents/2020/06/25/inaccessibility-of-emergency-services-number-112-on-24-june-2019/Inaccessibility%2Bof%2Bemergency%2Bservices%2Bnumber%2B112%2Bon%2B24%2BJune%2B2019.pdf
  8. https://www.inspectie-jenv.nl/actueel/nieuws/2020/07/02/veiligheidsregio%E2%80%99s-beter-voorbereid-op-crises-maar-nog-stappen-te-zetten
  9. https://www.rdi.nl/site/binaries/site-content/collections/documenten/2021/05/26/jaarbericht-2020/Jaarbericht%2BAgentschap%2BTelecom%2B2020.pdf
  10. https://ir.kpn.com/files/doc_financials/2019/ar/Integrated_Annual_Report_2019.pdf
  11. https://ir.kpn.com/news-and-events/events/event-details/2020/KPN-Annual-Report-2019/default.aspx
  12. https://ir.kpn.com/news-and-events/news/news-details/2020/Publication-of-KPNs-Integrated-Annual-Report-2019-02-24-2020/default.aspx
  13. https://wetten.overheid.nl/BWBR0009950/2020-12-21/0/
  14. https://wetten.overheid.nl/BWBR0032149
  15. https://wetten.overheid.nl/BWBR0043937/
  16. https://eur-lex.europa.eu/legal-content/EN/TXT/?qid=1657563539506&uri=CELEX%3A32018L1972
  17. https://www.etsi.org/deliver/etsi_ts/123100_123199/123167/14.05.00_60/ts_123167v140500p.pdf
  18. https://www.inspectie-jenv.nl/site/binaries/site-content/collections/documents/2019/08/22/plan-van-aanpak-crisiscommunicatie-112/Plan%2Bvan%2Baanpak%2Bcrisiscommunicatie%2B112%2Bdef%2Bpublieksversie.pdf