Summary

  • KDDI defines the main service-impact window as 61 hours and 25 minutes, from 01:35 JST on 2 July until 15:00 on 4 July 2022. It separately confirmed normal individual and corporate service at 15:36 on 5 July. Those clocks measure different things and should not be collapsed into one convenient outage number.
  • The initiating error was an incorrect route setting during planned maintenance of a nationwide transport-network router. Reverting that setting did not end the incident: a surge of location-registration retries congested VoLTE nodes and subscriber databases, produced inconsistent subscriber state, and made recovery traffic part of the failure.
  • Accountability follows practical control. KDDI controlled the maintenance gate, the rollback design, admission and congestion controls, observability, recovery sequence, customer information and redress. Regulators controlled serious-accident review and sector rules. Enterprise and public users could build fallbacks, but they did not operate the failed carrier control plane.
  • The effect reached beyond ordinary handset inconvenience. Official records connect the outage to emergency-call warnings, weather-observation gaps, logistics, connected vehicles, public administration, banking and transport. The selected evidence does not establish that the outage caused a death or any particular medical outcome.
  • Durable repair requires more than a list of completed actions. It requires replayable evidence that a route change and rollback remain safe under mass re-registration, that coupled VoLTE and subscriber-database congestion is detected early, that essential traffic has bounded alternatives, and that public status messages distinguish data, voice, emergency calling and final validation.

A recovery action turned a brief interruption into the real incident

The most useful way to read KDDI's July 2022 outage is to resist the hunt for a single dramatic moment. There was an initiating mistake, but there was also a recovery mechanism that amplified it. During maintenance of a router in KDDI's nationwide transport network in Tama, an incorrect route setting caused roughly fifteen minutes of communication interruption. KDDI reverted the setting. Devices and network equipment then sent large numbers of location-registration requests again.

VoLTE nodes became congested; subscriber databases used for authentication became congested; inconsistencies appeared in subscriber data; and controls intended to reduce load did not reduce it enough for a long time.

That sequence matters because "we rolled back" is often treated as the end of a change-management story. In a stateful network, rollback can be another high-impact change. The configuration may return to its previous value while millions of endpoints, sessions and databases remain in a new state. Devices retry. Timers expire together. Queues fill. Authentication systems receive queries arising from failure in another layer. Recovery commands compete with customer traffic. A control that is safe under normal demand may be unsafe when the whole population attempts to re-establish state.

KDDI says it ultimately identified VoLTE nodes generating unnecessary excessive signalling and separated them. That action helped restore both voice and data traffic to normal levels. The central accountability question is therefore not merely who entered an incorrect route. It is why a recoverable maintenance error could create a self-sustaining, nationwide congestion condition; which team could see that condition; what authority existed to isolate it; and what pre-tested recovery path could have stopped the cascade sooner.

This distinction protects the analysis from two weak conclusions. The first is individual blame: the assumption that an engineer's error fully explains a multiday event. The second is technological fatalism: the idea that a huge mobile network is inevitably too complex to recover predictably. Human error is foreseeable. Complexity is a design condition. Management is responsible for building approvals, simulation, observability, safe failure modes and recovery controls around both.

The public evidence is unusually rich, but it still has boundaries

The case can be reconstructed from company, regulator, public-agency and independent records. Each source answers a different question; none should be treated as a complete forensic image.

# Public record Use in this analysis
1 KDDI's incident, cause, prevention and refund record Primary chronology, population estimates, technical sequence, corrective measures and redress.
2 KDDI's detailed incident presentation Signalling flow, recovery stages, governance response and implementation dates.
3 KDDI's investor-relations outline Social-infrastructure duty and confirmation of the serious-accident report to MIC.
4 KDDI's 2022 integrated report Board-level synthesis of cause, scale, public-service effects, governance and regulation.
5 KDDI's 2023 integrated report Later claims that planned detection, restoration and work-control measures were completed.
6 KDDI's response to administrative guidance Procedure, congestion-control, early-recovery and communication improvements reported to MIC.
7 KDDI's financial-results presentation Board-facing remediation progress and critical-communications planning.
8 MIC's telecommunications-accident verification report Regulatory examination of expansion, duration, operational controls and public information.
9 Keitai Watch's summary of the MIC report Accessible independent cross-check of the regulator's principal findings and figures.
10 The Japan Meteorological Agency's AMeDAS impact notice Number of affected stations, peak data-collection gap and recovery challenge.
11 The JMA commissioner's later contract explanation Contract charge reduction and the boundary around monetary claims for missing data.
12 Nomura Research Institute's post-event survey Surveyed impact, user fallbacks and demand for emergency inter-carrier roaming.
13 Kofu Regional Fire Administration's 119 advisory Contemporaneous emergency instruction to use a landline or another carrier.
14 ITmedia's report on MIC intervention and public notice Ministerial criticism and the communication-accountability interval.
15 RCR Wireless's restoration report Independent English-language cross-check of restoration and sector impact.
16 The five-carrier JAPAN Roaming launch Later sector fallback, activation modes, device limits and emergency-call boundaries.
17 KDDI's affected-service status record Enterprise service families and the 5 July final validation point.
18 KDDI's warning about fraudulent refund messages The secondary trust risk created by large-scale redress communications.

The company records are strongest on what KDDI says happened inside its network and what it says it changed. MIC's report matters because it provides a regulator's examination rather than a corporate narrative alone. Public-agency records make downstream dependence measurable. Independent reporting helps establish what customers were being told in real time. Later corporate reporting is evidence that actions were claimed complete, not automatic proof that every control remained effective.

Several boundaries remain. The public record does not include the full change ticket, exact approval sequence, all alarms, per-node load histories, vendor roles or every recovery decision. It does not provide a final allocation of civil liability. It does not establish a unique number of affected human beings, because voice and data estimates overlap. It does not prove that any death or named medical event was caused by the outage. A disciplined article keeps those unknowns visible rather than filling them with confident invention.

Four clocks prevent a misleading outage timeline

KDDI's public material supports at least four clocks. The first begins at 01:35 on 2 July, when the service disruption started. The second ends at 15:00 on 4 July, when KDDI defines the main affected period as having ended after 61 hours and 25 minutes. The third reaches 15:36 on 5 July, when the company says it made a final confirmation that service use and traffic were normal for individual and corporate customers. The fourth continues through the serious-accident report, administrative guidance, customer refunds, corrective actions and later sector resilience work.

Calling the event a 61-hour outage is defensible if the speaker means KDDI's defined service-impact window. Saying the incident was not finally validated until the next day is also defensible. Calling it an 86-hour outage without explaining the difference blurs service restoration and assurance. The distinction is not cosmetic. It shows what a carrier must prove before using words such as "recovered", "restored" and "normal".

One validation gate concerns network traffic. Are call volumes, registration requests and database queries within expected ranges? A second concerns customer experience. Can users initiate and receive voice calls, use data, receive SMS and access dependent services? A third concerns the full population. Are there regions, device types, enterprise products or MVNO routes still failing? A fourth concerns stability. Does the network remain healthy after throttles are relaxed and temporarily isolated nodes are returned? A fifth concerns communication. Has the carrier told users which functions work and what they should do?

The July 5 final confirmation therefore belongs in the accountability record even if the measured disruption ended on July 4. It represents the time needed to move from apparent recovery to evidence-backed declaration. Good incident communication would publish these gates explicitly. Customers should not have to infer whether "recovery work completed" means that engineers changed a configuration, that data traffic is improving, that voice is usable, that emergency calls are reliable, or that the complete service population has been validated.

Rollback must be designed as a production workload

Change-management policy often asks whether a rollback plan exists. The KDDI event suggests a harder question: has the rollback been load-tested as a distinct production workload? Returning a route to its previous state may force millions of devices to recreate location, authentication and voice-service state. That behaviour can be more demanding than the original change.

A credible pre-change assessment would model at least three conditions. The first is a clean cutover in which sessions migrate as expected. The second is a partial failure in which some regions, nodes or databases accept the new path while others do not. The third is rollback after the population has already reacted to interruption. That third state contains synchronised retries, stale or inconsistent subscriber records, duplicate queries, delayed timers and operator actions taken under pressure.

The maintenance gate should therefore ask for numbers, not only procedures. What is the maximum expected registration rate after a fifteen-minute nationwide interruption? How much spare capacity do VoLTE nodes and subscriber databases have at that rate? Do retries include jitter, backoff and admission control, or do devices and network elements line up behind the same timer? Which queries are idempotent? How quickly can a node be quarantined? What happens when database correction itself produces more signalling? Which customer populations can be restored in stages?

Boards should not reduce this lesson to "prevent wrong configuration". Perfect input is not a realistic safety strategy. The stronger objective is to keep one wrong input from producing national common-mode failure. That means limiting blast radius, applying changes in stages, holding independent management access, rehearsing rollback under load, reserving capacity for recovery and defining who can stop or isolate the affected domain.

It also means preserving evidence. A post-incident reviewer should be able to reconstruct which version of the procedure was used, who approved each gate, what the router accepted, when alarms fired, how registration traffic changed, which nodes were isolated, which throttles were applied and why public declarations changed. Without that record, "human error" becomes a label that conceals the system around the human.

VoLTE and subscriber databases formed one recovery control plane

Mobile voice is not simply a call travelling through a switch. A device must establish network state, register its location and interact with authentication and subscriber systems before ordinary voice service can work. KDDI's account shows how failure in one layer can turn the recovery of another into a shared bottleneck.

The route interruption generated repeated registration signals. VoLTE nodes handled voice-related signalling. Subscriber databases received a surge of inquiries and developed inconsistent state. This coupling helps explain why a network can show partial data improvement while voice remains unreliable. It also explains why a generic "equipment fault" description would be inadequate. The failure involved the relationship between routing, endpoint behaviour, voice control and identity state.

Accountability should follow that relationship. The transport-network team may own the route change. A voice-platform team may own VoLTE nodes. A subscriber-platform team may own authentication databases. An operations centre may own incident command. Customer-service and communications teams may own public messages. If each team measures only its component, the organisation can miss the combined state that customers experience.

A better service model defines cross-domain indicators. Registration request rate is one. VoLTE success, call setup and incoming-call reachability are others. Subscriber-database query latency, error rate and consistency are essential. So are data-session success, SMS delivery, emergency-call completion and the experience of MVNO and enterprise products using the same underlying network. Those signals should be visible in one incident view with agreed thresholds and owners.

Admission control also needs ethical priority. During a mass retry event, not every request has equal consequence. Emergency calling, incident command, public warning, transport safety and essential telemetry may justify reserved capacity or an alternate path. The engineering answer is constrained by standards, device behaviour, interconnection and regulation; it cannot be improvised during an outage. Priority and fallback must be designed before the congestion begins.

Database consistency deserves special attention. Correcting inconsistent subscriber records may restore some users while creating more queries or causing other devices to retry. Recovery tooling should therefore be tested on a realistic copy of failure-state distributions, not only on clean laboratory data. It should report how many records are corrected, how many remain uncertain, what secondary signalling it produces and how operators can reverse a harmful batch.

The goal is not to eliminate every dependency between network functions. It is to make dependencies observable, bounded and recoverable. A nationwide carrier should know which control-plane components can fail together, what traffic they generate during repair and which independent management channels remain when ordinary customer and staff communications are impaired.

Scale must be reported without manufacturing a bigger number

KDDI estimated about 22.78 million affected voice users and at least 7.65 million affected data users in its own operations. With Okinawa Cellular included, it reported about 23.16 million voice users and at least 7.75 million data users. These figures are not four independent populations. A person or line can be counted in both voice and data estimates, and the methodology compares activity during the incident with a normal reference period.

Early public reporting also used potential-impact figures approaching 39 million connections. That number describes a different boundary from the later affected-use estimates. Both can be relevant, but only if labelled. Potential exposure asks how many connections belonged to the service population that could have been affected. Estimated actual effect asks how observed calls or registration requests differed from normal. A unique-person estimate would require de-duplication that these public figures do not supply.

The temptation to add voice and data totals should be resisted. A larger headline does not strengthen accountability; it weakens it by making a category error. The 61-hour duration, nationwide reach and public-service effects already establish materiality. Accurate denominators allow customers, regulators and engineers to compare incidents and judge whether repair worked.

A handset outage became a public-infrastructure outage

KDDI's integrated report identifies impacts in logistics, automobiles, administrative services, banking and transportation. Delivery status updates and communication with drivers were affected. Connected-car services depended on the network. Weather-data collection and water meters represented public and municipal dependencies. Some off-premises ATMs used the carrier. Airport wireless transceivers and bus IC-card systems appeared in the company's impact map.

The Japan Meteorological Agency provides the clearest quantified downstream example. Of 1,284 AMeDAS observation stations, 802 were intermittently affected during the period, and roughly 550 could not deliver data simultaneously at the peak. Some observations could later be collected from storage at the local equipment, but recovery required field and quality work, and long-duration gaps created a risk that not every value could be obtained.

This is a different kind of telecom harm from a person unable to stream video. Weather observations enter forecasts, warnings, climate records and public decision-making. The immediate service may have local storage, yet delayed collection changes timeliness. A transport or banking dependency may have a manual fallback, but the fallback costs staff time, delays transactions or reduces capacity. A connected vehicle may remain physically safe while remote functions fail. One carrier incident therefore creates many different impact clocks.

Dependency accountability has two owners. The carrier must understand how its connections are used and which classes require additional protection or communication. The customer must understand whether a mobile link is a convenience, a primary operating path or a safety-critical single point of failure. A contract can assign service levels and credits, but it does not automatically create technical diversity.

Enterprises should ask whether a second SIM or carrier is genuinely independent. Two products can share radio infrastructure, transport, authentication, power, sites or operational tooling. A fixed line may offer better diversity, but it can share a facility or backhaul. Satellite service may help selected functions, but it introduces device, capacity and sky-view constraints. The objective is not to buy two logos; it is to map common failure domains.

KDDI likewise needs a machine-service inventory. Which connections carry emergency or safety-related telemetry? Which devices retry aggressively? Which applications can buffer locally? Which services require incoming calls, SMS or stable identity state rather than simple data reachability? Which enterprise customers cannot receive an ordinary consumer notice during the outage because their own communication channel uses the same network?

The AMeDAS record also shows why evidence must continue after connectivity returns. Restoration of the link does not prove restoration of the data. Operators and customers need reconciliation: which observations or transactions were delayed, recovered, duplicated, rejected or permanently unavailable? Without that ledger, a green network dashboard can conceal unfinished downstream recovery.

Emergency access is the hard boundary, not a rhetorical flourish

A regional fire administration warned that KDDI users might not be able to call 119 and told people to use a landline or another carrier. The significance is immediate. Mobile connectivity is a safety service when a person needs police, fire, ambulance or coast-guard assistance. Advising an alternate path is necessary, but it assumes that the user has one, understands the instruction and can act under stress.

The public evidence used here does not establish that the outage caused a death or a specific medical outcome. Responsible analysis should not imply one. The absence of proven causal harm, however, does not remove the duty to protect emergency reachability. Safety controls are judged partly by the credible consequence they are designed to prevent, not only by the worst outcome that can later be proved.

Emergency continuity has several layers. The network needs a protected path for call setup and location information. The receiving authority needs to accept and, where necessary, return a call. Devices need compatible behaviour. Users need clear instructions. MVNO customers need to know whether their service inherits the same failure. Public authorities and carriers need activation criteria for alternatives. Every layer creates evidence that can be tested before an incident.

The NRI survey illustrates public expectations without turning a survey into a national census. In a screening sample of 28,984 people, 20.5 per cent reported being affected. Among survey respondents, 72.4 per cent said inter-carrier roaming was necessary or somewhat necessary. Emergency calling appeared in the top five desired roaming functions for 88.2 per cent. These figures describe the sampled population and methodology; they do not prove exactly how every resident would answer.

They do show that resilience cannot rely entirely on personal improvisation. Wi-Fi calling, messaging applications, public phones, landlines and dual SIM can help, but access varies. A person away from home may have no fixed line. A power failure may remove Wi-Fi. A device may not support the relevant mode. An emergency caller may not know how to select another network. Sector design must assume uneven resources.

The accountability test for KDDI is therefore not whether it could promise perfect availability. It is whether emergency failure modes were identified, bounded, communicated and backed by interoperable alternatives. The corresponding test for government is whether rules, standards and exercises create a usable fallback across carriers rather than a policy that exists only on paper.

Public communication should expose service gates, not optimism

During a prolonged recovery, language can transfer risk. If a carrier says recovery work is complete while voice calls remain difficult, users may stop seeking alternatives. If it says the network is gradually recovering without distinguishing data and voice, businesses may resume processes that depend on incoming calls or SMS. If an enterprise status page is reachable only through the failed network, accurate text is still operationally unavailable.

The ministerial criticism reported during the incident focused on limited public evidence public notice. KDDI later listed improvement of information disclosure and timely, appropriate announcements among its corrective actions. That is not a public-relations add-on. Communication is a control that helps customers reduce harm while engineers repair the network.

A useful status message should identify the affected service, geography and population; the current customer experience; emergency-call status; enterprise and MVNO implications; the last verified time; the next update time; available fallbacks; and what "recovery" means at that stage. It should separate engineering action from customer validation. It should avoid asserting a cause before evidence supports it.

Updates should also be distributed through independent channels. Web pages, television, radio, other carriers, public agencies, in-store notices and service-specific contacts may all be necessary when SMS and mobile data are unreliable. Enterprise customers need an out-of-band contact that does not depend on the same carrier account. Emergency authorities need a direct operational channel rather than learning from a consumer page.

Large-scale redress creates a second communication risk. KDDI warned about fraudulent email and SMS messages exploiting its refund programme. A company can unintentionally create a phishing theme by telling tens of millions of people to expect money or account contact. Redress messages should therefore be recognisable, avoid unnecessary links, state what information will never be requested and provide a separately verified route for checking the credit.

Refunds made redress visible without pricing every harm

KDDI described two forms of refund. A terms-based refund applied to 2.71 million KDDI customers and 70,000 Okinawa Cellular customers who could not use all communications for more than twenty-four consecutive hours or were in an equivalent position under eligible voice-only contracts. It represented two days of the relevant basic charge.

An apology refund had a much larger population: 35.89 million KDDI customers and 660,000 Okinawa Cellular customers with eligible smartphone, feature-phone and Home Plus Phone services. KDDI set the amount at 200 yen excluding tax. Because povo2.0 had a zero-yen base fee, eligible users received a one-gigabyte, three-day data topping instead.

These decisions answer practical questions. Who had to apply? How would customers who changed plans or cancelled receive the credit? How would corporate or fixed-service users be contacted? What alternative made sense for a zero-base-price product? KDDI's published process reduced the burden by making much of the redress automatic.

The amounts do not measure every harm. A missed delivery, inaccessible ATM, unavailable business call, delayed weather observation and failed personal conversation do not have the same value. An emergency risk cannot be responsibly reduced to a pro-rated subscription fee. Nor does a customer credit resolve public-sector recovery costs or contractual disputes.

The Japan Meteorological Agency later explained a useful boundary. Under its contract, it would reduce line charges by about 800,000 yen for unavailable service. It did not pursue damages for missing observations because it could not assign those data a monetary value in the way required for that claim. That does not mean the observations were worthless. It means contract credit, legally recoverable damages and public value are different categories.

Boards should examine redress as part of incident design, not as a late commercial decision. Systems should be able to identify affected services and periods, calculate eligibility, prevent double payment, handle product changes and create an appeal route. Fraud risk must be anticipated. Contract terms should not be the only lens; voluntary redress may recognise trust and inconvenience that a strict service formula misses.

Regulation turned an internal failure into a public control record

KDDI submitted a serious-accident report under Article 28 of the Telecommunications Business Act on 28 July. MIC issued a reprimand and written administrative guidance on 3 August. The regulator's verification process examined why the event expanded, why recovery took so long, how work procedures and congestion controls operated, and how information reached customers.

This public review matters because a carrier controls most primary evidence. It owns the telemetry, tickets, diagrams, staff interviews and recovery logs. Customers see symptoms. Independent reporting sees statements and selected effects. A regulator can demand a more complete record and compare it with statutory duties and sector practice.

KDDI's governance response put the president in charge of a council for improving infrastructure and customer relations, with all executives involved. Working groups covered work quality, operations, equipment and customer relations. That structure recognises that the failure crossed technical and organisational boundaries.

It can still become ceremonial. Executive ownership is meaningful only if it resolves conflicts. A maintenance team may want speed; safety review may want staged rollout. Network operations may want to protect capacity; customer teams need frequent information. Product groups may resist expensive diversity. Procurement may accept a supplier's assurance without a joint recovery test. The council must decide, fund and track those trade-offs.

KDDI reported reviewing procedure management and approval, risk assessment and work-suppression criteria. It developed more detailed congestion detection, reviewed congestion-control design, revised recovery procedures, built tools to relieve VoLTE congestion and improved public communication. Later reporting said planned measures for the fiscal year had been completed.

Completion is the start of assurance, not the end. A tool can be installed but not used. A dashboard can show congestion after customer harm has already spread. A one-touch recovery control can automate the wrong action. A new procedure can drift out of date. Regulatory closure should therefore ask for exercise results, incident metrics, independent tests and evidence that operators can use the controls under realistic pressure.

The event also raises a sector question. When mobile networks support public safety and essential machine connections, how much resilience should be left to each carrier's private design, and which capabilities require common rules? Serious-accident reporting creates learning only if findings change shared assumptions about emergency access, roaming, customer notice and critical-service dependencies.

Verifiable repair needs hostile but safe rehearsal

KDDI's later integrated report describes congestion detection that visualises service impact on a dashboard and restoration measures that operators can trigger across multiple exchanges. It also describes measures intended to prevent human error and a strengthened quality organisation. Those controls address the observed failure modes. The remaining question is whether they work under the conditions that made July 2022 difficult.

A credible rehearsal would begin with a route-setting error in a contained environment that models nationwide endpoint behaviour. It would allow some connections to fail, then trigger a rollback and mass re-registration. The test would deliberately create VoLTE-node congestion, subscriber-database query pressure and inconsistent state. Operators would have to identify the coupled condition, apply admission controls, isolate harmful nodes, correct records, protect emergency traffic and communicate service gates.

The exercise should produce a durable evidence set. It should record change approvals, configuration versions, synthetic endpoint populations, retry distributions, node capacity, database latency, alert times, operator decisions, isolation time, customer-function tests and public-message drafts. Results should show not only that recovery succeeded, but which thresholds failed and what residual risk remains.

Canary deployment is another control. A nationwide route change should not become nationwide merely because the configuration entity has national scope. The organisation should identify a representative but bounded population, observe service and control-plane indicators, and require an explicit gate before expansion. The rollback should be canaried too where architecture permits.

Independent management access is essential. A carrier responding to a network outage cannot assume its staff will retain ordinary mobile voice, messaging, authentication or remote access. Incident command needs diverse communications, pre-authorised access, offline procedures and reliable contact with suppliers, government and emergency authorities. A recovery plan that depends on the failed service contains its own common-mode weakness.

Enterprise drills should participate. Weather, transport, banking and logistics users need to test local buffering, reconciliation and alternate connectivity. Their results help KDDI understand which service states matter. "Data restored" may be limited public evidence if delayed observations must be reconciled, an ATM needs secure session state, or a vehicle requires an inbound control channel.

Assurance should include near misses and smaller incidents. Metrics such as unsafe change caught before release, rollback signalling peaks, time to detect cross-domain congestion, emergency-call success under isolation, status-message accuracy and unresolved subscriber records provide a continuing view. Publishing selected aggregate measures would let customers and regulators judge improvement without exposing sensitive network detail.

The strongest repair statement is therefore not "all actions completed". It is: the same class of failure was reproduced safely; detection happened within a defined time; the blast radius remained below a defined bound; emergency functions retained an alternative; operators recovered within an objective; downstream data reconciled; and an independent reviewer could verify the evidence.

JAPAN Roaming is meaningful follow-through with explicit limits

In March 2026, KDDI, NTT DOCOMO, Okinawa Cellular, SoftBank and Rakuten Mobile announced JAPAN Roaming, scheduled to launch on 1 April. The service allows temporary use of another carrier's 4G LTE network when a major disaster or outage disrupts the contracted provider. It is a substantial sector response to the broader problem exposed by major mobile failures.

The design includes two modes. Full roaming can provide voice, data at a stated limited speed and SMS. Emergency Calls Only permits outgoing calls to 110, 119 and 118 but does not provide ordinary data or SMS. The release explicitly says callback is unavailable in that emergency-only mode. Device support, activation, coverage and user settings also matter.

Those boundaries belong in the safety case. An emergency centre may need to call a person back after a disconnected or incomplete report. A user may see "no service" while the instructions ask them to select another network and retry. Older devices may not support every feature. MVNO treatment differs by mode. A stressed caller may not know which carrier is the contracted one or how to reverse manual network selection afterwards.

The service should therefore be evaluated through exercises involving carriers, device makers, MVNOs, police, fire, ambulance, coast guard and representative users. Tests should include congestion on the donor network, simultaneous outages, regional disasters, power loss, inaccessible status pages and calls that require callback. Activation authority and customer notification must be fast enough for the service to matter.

It would overstate the evidence to say the KDDI incident alone caused JAPAN Roaming. Japan has a longer history of disaster and outage planning, and the final service reflects years of multi-party work. The 2022 event clearly sharpened public and policy attention to emergency alternatives, but causal credit belongs to a wider record.

It would also be wrong to treat roaming as permission for the primary carrier to reduce resilience. Donor capacity is finite. Major disasters can affect several networks and shared sites. Roaming is a safety net, not an ordinary substitute for safe maintenance, diverse control planes, adequate capacity and tested recovery inside each carrier.

The accountable claim is narrower and stronger: by 2026, Japan's major carriers had created an interoperable fallback that addresses a class of harm visible in 2022, while openly documenting important limits. Future audits should measure activation time, call success, callback handling, compatible devices, MVNO coverage and customer comprehension.

Responsibility should follow capability, not the first visible mistake

The incorrect route setting is part of the causal record. It does not settle the allocation of responsibility. An individual operator may control a command, but the institution controls who may run it, which document is authoritative, how risk is assessed, whether the change is staged, what automation validates it, how rollback behaves and when independent approval is required.

KDDI therefore carries the largest operational accountability. It had practical control over maintenance, the national network architecture, congestion design, telemetry, incident command, customer communications and redress. That does not mean KDDI caused every downstream choice or that it owes every alleged loss. It means the company held the capabilities most directly able to prevent, limit, detect, explain and repair the failure.

Okinawa Cellular had customer and service responsibilities within its operating scope and participated in redress. The selected public material does not justify inventing a percentage division of the technical cause between the companies.

Equipment suppliers and contractors would be responsible for product behaviour, support or validation within their actual roles. The selected public evidence does not identify a supplier defect as the root cause. Naming one would convert a gap into an accusation.

MIC and public authorities controlled serious-accident review, administrative guidance and sector policy. They did not operate KDDI's routes, VoLTE nodes or subscriber databases. Their accountability concerns whether oversight surfaced the right failure modes, whether guidance was specific and verifiable, and whether emergency alternatives progressed.

Enterprise and public-sector users controlled their dependency maps, local fallback and downstream reconciliation. They could choose diverse communications in some contexts. Their options were constrained by cost, shared infrastructure, standards and the carrier services available. Customer resilience does not erase carrier duty.

Individuals had the least practical control. Some could use Wi-Fi, a landline, a public phone, another SIM or another person's device. Others could not. It would be perverse to make personal preparedness the primary answer to a national control-plane failure. Clear fallback advice is useful; it is not a transfer of architectural responsibility to the user.

The mobile-carrier sector now shares responsibility for emergency-roaming interoperability. Each carrier must maintain its primary network and contribute a safe donor network when the agreed conditions apply. Government must ensure that emergency-calling rules, device behaviour and public instructions match real operating conditions.

This capability-based allocation avoids two extremes. It does not absolve the person who ignored a procedure, if evidence shows that occurred. It also does not let the board declare the event solved by retraining an operator. Accountability rises with the ability to change the system around the error.

The questions that should remain on the board agenda

A board does not need to configure VoLTE nodes to govern this risk. It needs a concise set of evidence questions that connect technical behaviour to public duty.

First, what is the largest failure domain for a routine maintenance change? The answer should identify customer population, geography, services, enterprise dependencies and emergency functions. It should show which change can still reach nationwide scope and why.

Second, has rollback been tested under realistic endpoint response? A slide saying "rollback available" is limited public evidence. The board should see the assumed registration rate, tested rate, spare capacity, admission controls and time to safe service.

Third, can the organisation detect cross-domain congestion before customers report it? Evidence should connect routing, VoLTE, subscriber databases, data sessions, SMS, emergency calls and enterprise products.

Fourth, which recovery functions depend on the affected network? Staff communications, multifactor authentication, remote access, vendor contact and public status publishing need diverse paths.

Fifth, how does incident command distinguish restoration from validation? Definitions should exist for engineering work complete, partial data service, voice recovery, emergency-call availability, regional recovery, enterprise recovery and final normality.

Sixth, which essential customers rely on single-carrier machine links? The carrier and customer should know local buffering, reconciliation, manual operation and alternate connectivity. Contract credits should not substitute for technical plans.

Seventh, what evidence supports completed remediation? Exercise results, control use, alert times, canary coverage, independent review and residual risk are stronger than action counts.

Eighth, how quickly can emergency roaming be activated, and what does it not provide? Callback, device, MVNO, capacity and simultaneous-failure limits should be visible.

Ninth, can redress be calculated and communicated without inviting fraud? Automated credits, authenticated notices, no-link guidance, product changes and appeals should be rehearsed.

Tenth, what will be disclosed after the next major event? A credible template should preserve uncertainty, publish function-specific status and provide enough technical explanation for independent learning.

These questions make the KDDI event more than a history lesson. They turn a public incident into a repeatable governance test for every carrier, cloud platform and identity system whose recovery traffic can become its next outage.

Recovery is accountable only when its safety can be shown

KDDI's July 2022 outage did not prove that large networks cannot be reliable. It proved that reliability cannot be inferred from stable configuration alone. A maintenance error, a rollback, device retries, VoLTE nodes and subscriber databases interacted to create a new production state that the organisation struggled to reduce.

The public response contains real signs of accountability: a serious-accident report, regulatory review, customer refunds, cross-functional governance, published corrective measures and later sector roaming. Those actions are stronger when their limits are stated. Refunds did not price every harm. Completed measures did not certify permanent effectiveness. Emergency-only roaming does not support callback. Public population figures do not produce a unique-person total.

The durable lesson is that responsibility follows control over the recovery path. The accountable operator can show that changes are bounded, rollback is load-tested, coupled congestion is visible, emergency access has an alternative, status words map to service evidence, downstream data is reconciled and remediation survives hostile but safe rehearsal.

That is a higher standard than avoiding the same incorrect route. It is also the standard appropriate to infrastructure that people use not merely to communicate, but to call for help, move goods, observe weather, access money and operate public services.