Summary
- The incident began at about 01:35 JST on 2 July 2022 during maintenance of a nationwide transport core router at KDDI's Tama network centre. A wrong or obsolete procedure omitted a VoLTE routing change, leaving an intermittent response failure that triggered repeated location-registration requests.
- The retry load spread through distributed VoLTE call control and subscriber databases. Six call-control nodes later restarted from backups written while internal state was inconsistent, so reset did not reliably return them to a clean condition.
- Recovery required more than reversing the route change. KDDI used flow controls, call-control resets, PGW isolation and session resets, subscriber-database load redistribution, and ultimately isolation of six persistently abnormal VoLTE nodes.
- The official affected-service window was 61 hours 25 minutes, from 01:35 on 2 July to 15:00 on 4 July. KDDI did not confirm all telecommunications networks fully operational until 15:36 on 5 July, about 86 hours after the start. Those are different operational clocks.
- The Ministry of Internal Affairs and Communications treated the event as a serious accident and identified failures in procedure control, congestion design, high-load recovery testing, incident exercises, and public communication. KDDI's refunds and reported recurrence measures were operator responses, not a regulator-imposed monetary fine or a court finding.
The incident has more than one clock
KDDI places the start of the communications failure at about 01:35 Japan Standard Time on Saturday, 2 July 2022. Its consolidated English incident page gives an affected period ending at 15:00 on Monday, 4 July, a total of 61 hours 25 minutes. The Ministry of Internal Affairs and Communications, or MIC, used the same affected-time window in its dedicated verification report.
Kyodo reported a later milestone: KDDI confirmed that all telecommunications networks were fully operational at 15:36 on Tuesday, 5 July. That was about 86 hours after the incident began. The longer interval should not replace the 61-hour-25-minute measure, and the shorter interval should not erase the additional testing and normalization work. They answer different questions.
One clock describes the period KDDI and the regulator assigned to affected services. Another ends when the operator confirmed normal operation across all networks. Between them sit the removal of flow restrictions, checks of network stability, observation of voice success, correction of inconsistent state, and the decision that the system could again carry normal demand. A report that calls any one of those steps simply "restored" loses the evidence needed to judge recovery.
The impact estimates also belong to source-specific snapshots. The MIC report estimated approximately 23.16 million people affected for voice and at least 7.75 million for data, a cumulative estimate of at least 30.91 million. KDDI's later consolidated page reports approximately 22.78 million for voice and at least 7.65 million for data. Reuters, reporting while the detailed cause was still being investigated, cited an early potential exposure of up to 39.15 million users, including 260,000 corporate customers.
Those figures cannot be averaged or added. The early maximum described potential exposure during an unfolding incident. The MIC and KDDI figures came from different official snapshots and calculations. Voice and data populations can overlap. Contracts, connections, people, and service attempts are not interchangeable units. Preserving each definition is not statistical caution for its own sake; it is what keeps a large incident from becoming a number assembled by the writer rather than one established by the evidence.
An omitted route change created a partial failure
The technical sequence began during maintenance of a nationwide transport core router at the Tama network centre. The MIC found that the work used a wrong or obsolete procedure. The procedure did not direct the necessary routing change for VoLTE traffic, leaving traffic pointed toward a path that was no longer operating as intended.
The resulting failure was intermittent rather than cleanly total. The MIC reconstructed a condition in which responses associated with the affected path failed about half the time. That distinction mattered. A device or network function that receives no successful response may declare a hard failure and move to a known alternative. A system that succeeds sometimes and fails sometimes can remain just healthy enough to keep trying the same path.
Mobile voice over LTE depends on location registration. A device must establish that it is an authorized subscriber and register where it is reachable before normal incoming and outgoing voice calls can complete. When acknowledgements did not return, devices and network functions treated registration as incomplete and retried.
Retries are individually rational. They become dangerous when millions of endpoints and distributed control functions make the same decision at once. Each lost response generated another request. Each additional request consumed more call-control and subscriber-database capacity. As capacity fell, more requests failed, which produced more retries. The network began amplifying the initiating fault.
The MIC estimated that the Tama VoLTE call-control functions experienced roughly seven times normal load. Other VoLTE sites experienced about 6.5 times normal load as the architecture distributed work and retries across the country. A local maintenance error had entered a nationally coupled control system.
This is the first accountability boundary. The maintenance mistake explains why abnormal traffic began. It does not by itself explain why the failure spread nationwide or lasted for days. That requires examining how the running topology distributed signalling, how retry logic behaved under partial response loss, and whether congestion protection recognized this particular state as abnormal.
Congestion crossed from voice control into subscriber state
The VoLTE nodes did not operate alone. Location registration and voice-session setup depended on subscriber databases used for authentication, policy, charging, and location or service state. As VoLTE retries increased, requests toward those databases increased too.
The MIC describes a reinforcing loop. Subscriber-database congestion produced error responses. Those errors led devices and network functions to send more requests. More requests drove additional congestion. Some device behavior also meant that failed voice-session establishment affected the device's registered state, so data service could become unavailable intermittently even though the data-specific subscriber database was not the original point of congestion.
This boundary matters for public description. It would be inaccurate to say every mobile data service was continuously down across Japan. The evidence supports a nationwide communications failure with voice, SMS, and device-dependent or flow-control-related data effects. Different devices and services could show different symptoms at the same time.
Congestion also damaged the consistency of control state. The MIC found that repeated authentication-session changes under high load were not synchronized correctly inside some VoLTE call-control nodes. Six nodes wrote periodic backups while their internal state was inconsistent. When operators reset the call-control functions using the latest backups, those six nodes returned with the abnormal state still present.
Reset was therefore not equivalent to recovery. It restarted equipment, but the state loaded at startup continued to produce authentication failures and repeated registration requests. The action that normally clears a fault preserved part of the fault because the recovery artifact itself reflected the congested system.
A further inconsistency developed among the VoLTE functions, subscriber databases, and packet gateways, or PGWs. Session information that should have been removed was not always deleted while the database was congested. Later sessions could then conflict with state that remained in the system. Correcting that mismatch took time and had to respect the capacity of equipment that was already under pressure.
This is the outage's strongest running-code lesson. A backup file is a record, but it is only as trustworthy as the state captured in it. A documented reset procedure is an intention, but the live result depends on the data loaded, the traffic arriving, and the interactions among systems. In a complex incident, operators need proof that recovery has moved the network to a clean state, not merely proof that the command completed.
Traffic control was necessary but not sufficient
KDDI's response involved several different forms of intervention. The MIC lists controls on connection-request traffic and VoLTE traffic, resets of call-control functions, isolation of some PGWs to reduce subscriber-database load, resets of PGW sessions to correct inconsistent session information, and changes in which subscriber-database resources received PGW traffic.
These actions addressed different parts of the loop. Flow control reduced new demand. PGW isolation and session resets changed the state presented to subscriber systems. Load redistribution sought to keep one database path from remaining the limiting point. Call-control resets attempted to clear abnormal processing. No single action can fairly stand for the whole recovery.
The six VoLTE call-control nodes with persistent abnormal state were especially difficult. The MIC found that an attempted blocking process did not complete normally because the nodes were already abnormal. They continued generating excessive signals after earlier steps had reduced other sources of load. The eventual isolation of those six nodes stopped the persistent retry source, after which congestion subsided.
The sequence exposes a common failure in restoration language. An operator may finish the physical or configuration work it planned, while the running network remains congested. It may restore most data paths while voice calls still face restrictions. It may reach normal call-success rates while continuing to inspect every relevant system because long-running high load can leave hidden state behind.
The MIC later criticized the distinction KDDI communicated between completion of recovery work and full service resumption. KDDI intended completion of work to be followed by removal of flow controls and confirmation of normality. Some users and news organizations understood "recovery work completed" to mean service had recovered. In reality, congestion and traffic restrictions continued.
That was not only a communications problem. It was a measurement problem exposed through language. If the status system cannot state which technical and service conditions have passed, public updates will reach for vague milestones. A stronger incident model publishes separate evidence for the fault input being removed, retry growth being contained, congestion clearing, session state agreeing, flow controls being lifted, voice and data success stabilizing, and end-to-end dependent services being checked.
The outage reached beyond handsets
KDDI's own account lists effects on delivery updates and driver communications, connected-car services, weather-data collection, water meters, bank ATMs, airport wireless communications, and bus payment systems. Reuters independently reported disruption involving weather data, parcel delivery, banking, and connected vehicles.
These examples show why mobile-core continuity is a shared infrastructure concern. A mobile connection may be a control or reporting path inside another service. The visible failure for one user is a phone that cannot place a call. For a logistics operator it can be the loss of status updates. For a public body it can be missing observations. For a transport or banking service it can be the inability of a remote endpoint to complete a routine transaction.
The examples must retain a limit: they do not establish that every listed service failed nationwide for the entire outage. They demonstrate dependency reach, not a universal service-by-service clock.
Emergency calling carries the sharpest continuity consequence. Reuters reported the communications minister's concern that fire and emergency calls used to protect life and property had been hindered. Kyodo later reported that some users could not dial Japan's 110 or 119 emergency numbers for an extended period. The source set does not provide a count of failed emergency-call attempts, unique emergencies, injuries, or deaths.
That absence should not be filled with speculation. The supported finding is already serious: access to emergency numbers was impaired during a long nationwide mobile disruption. The accountability question is whether the operator can show that emergency communication has a survivable path when voice registration, subscriber authentication, or ordinary traffic control is failing.
The regulatory record identifies control failures, not every liability
The MIC treated the outage as a serious accident and convened a dedicated technical verification. Its report did not reduce the cause to one human error. It identified deficiencies in procedure selection and approval, service-normality checks, risk assessment, congestion-control design, evaluation of nationwide propagation, high-load reset testing, complex recovery procedures, incident exercises, and customer communication.
That layered finding is important. A wrong procedure can initiate an event, but national resilience depends on controls that assume procedures and people will sometimes fail. Partial-response states must be recognized. Retry behavior must be bounded. Distributed architecture needs containment zones. Backups and resets must be tested under abnormal load. Operations teams need a practiced route out of states that ordinary automation does not resolve.
KDDI reported changes including stronger procedure and approval controls, better congestion detection, review of congestion-control design, revised recovery procedures, recovery tooling, and improved customer information. Those responses correspond to the documented failure paths. The available public record does not provide a current, independent, fleet-wide test of their effectiveness in 2026.
KDDI also announced customer compensation. Its consolidated page separates terms-based refunds for customers who could not use all communications for more than 24 consecutive hours or an equivalent condition from a 200-yen apology refund for a much larger customer group. Kyodo reported an expected cost of about 7.3 billion yen and a voluntary salary reduction by KDDI's president.
Those are operator compensation and governance responses. They are not a regulator-imposed monetary fine, a criminal disposition, or a judicial finding of liability. The public record supports the serious-accident classification, the MIC verification findings, reported corrective measures, and the compensation program. It does not support turning those outcomes into a legal judgment that the sources do not describe.
Continuity evidence must describe the running service
The incident suggests a practical evidence hierarchy. At the first level is intended state: the approved procedure, expected route, configured congestion controls, recovery plan, and customer-notification policy. Those records matter, but the outage shows that they cannot establish continuity alone.
The next level is observed technical state. Operators need to know which path VoLTE signalling is actually using, whether acknowledgements are returning, how retries are distributed, which call-control nodes are synchronized, whether subscriber and PGW session state agrees, and whether a reset loaded a clean backup. These measurements show what the network is doing rather than what its documents say it should do.
The highest level is service completion. Successful node health does not prove that users can register, place and receive calls, send messages, use data, or reach an emergency centre. Recovery needs end-to-end checks across representative devices, regions, services, and dependent organisations. The public milestone should follow those outcomes.
Regulatory reports preserve a vital ledger of what happened. They fix terminology, figures, chronology, and corrective expectations. But the report cannot carry traffic, stop retries, reconcile sessions, or make an emergency call complete. Its operational value lies in the tests it creates for the running network.
That is why the decisive question after KDDI's 2022 outage is not whether the original route was corrected. It is whether present controls can prove that a similar intermittent failure stays local, retry demand remains bounded, recovery artifacts remain trustworthy, operators retain visibility, and public status language maps to observable customer outcomes.
What the public evidence does not establish
The MIC report provides a detailed reconstruction, but the public sources do not expose every internal log, each command in the recovery sequence, or every individual decision owner. They do not establish a complete per-device impact census. They do not quantify emergency-call failures or downstream economic loss.
The official impact estimates also changed between snapshots. That does not make the sources unusable. It means each figure needs its owner and method. The early 39.15 million potential-exposure figure belongs to contemporaneous reporting. The MIC's 23.16 million voice and 7.75 million data estimates belong to its verification report. KDDI's later 22.78 million and 7.65 million figures belong to its consolidated customer page.
The evidence also cannot establish current effectiveness from a list of measures implemented or planned in 2022. A tool may have been deployed without covering every relevant node. A procedure may have changed without being tested against intermittent loss and nationwide retry load. A recovery exercise may pass under laboratory conditions that do not reproduce inconsistent subscriber state.
These limits do not weaken the accountability case. They keep it tied to what can be demonstrated. The incident began with a concrete maintenance and routing failure, spread through observable retry and congestion mechanisms, persisted through inconsistent state, affected critical dependencies, and required a multi-stage recovery. The responsible conclusion is about controls and evidence, not invented harm or personal blame.
Sources
- KDDI, The July 2 Communication Failure and Our Response
- Japan MIC Telecommunications Accident Verification Council, verification report on the 2 July 2022 KDDI and Okinawa Cellular serious accident
- Reuters via Euronews, KDDI aims to restore service Sunday after 40 million users affected
- Kyodo News, KDDI to pay damages to 36 mil. people affected by network outage
Member Briefing
Deeper Profile Context
Sign in with the right membership level to unlock the full briefing and source notes.
Only for Strategic Circle
Strategic Circle
Open to all readers. Unlock profile briefings after joining and signing in.
Join Strategic CircleOnly for Leadership Alliance
Leadership Alliance
For qualified IP-asset owners and management; sign in to unlock alliance briefings.
Join Leadership Alliance
