Summary
- NTT Docomo was migrating an IoT subscriber and location-information server from old equipment to new equipment on 14 October 2021. After a problem involving some overseas-roaming IoT behavior was found, the operator switched the population back. Its technical report says a procedural misunderstanding with a contractor caused a large population of devices to return at once and emit a mass of location-registration signals. [1]-[5]
- The surge did not remain within an IoT service. Docomo and Japan's Ministry of Internal Affairs and Communications said IoT and ordinary mobile users shared signalling-switch processing resources. The registration load exhausted those resources, created congestion between subscriber or location servers and signalling switches, and propagated across the nationwide network. [3][5][7][8]
- Public impact measurements describe different conditions and must not be added into one count of unique people. The unable-to-use period ran from 17:37 to 19:57 Japan Standard Time, lasting two hours and twenty minutes, with approximately one million users estimated to have been affected. A difficult-to-use condition ran from 16:54 on 14 October to 22:00 on 15 October, lasting 29 hours and six minutes; the operator estimated approximately 4.6 million voice users and at least 8.3 million data users were affected by that condition. [3][7]
- Recovery was staged. The operator controlled 4G location registration, adjusted the flow of IoT registrations, restored 5G and 4G before 3G, and continued service-specific work after the most acute unable-to-use period ended. That sequence shows why lifting one restriction or returning one server does not prove that customer reachability has recovered. [1][3][5][7]
- Japan's regulator treated the event as a serious accident and required action on migration preparation, contractor coordination, isolation between IoT traffic and voice or other communications, emergency-call communication, and industry learning. The classification establishes a formal public-interest threshold, not by itself negligence or civil liability. [6]-[8]
- Docomo's announced remediation included comparing old and new specifications, adding overseas-roaming tests, aligning cutover and switchback procedures, setting decision deadlines, supporting IoT-only registration control, separating registration-processing resources, and exercising network-control procedures. These are meaningful control commitments, but public documents alone do not prove complete deployment or continuing effectiveness. [4][5]
- Independent downstream evidence from IIJ records effects on voice, data, M2M, and IoT services using Docomo's network. It confirms dependency propagation beyond Docomo's retail notice without revealing every private path or customer effect. [9]
- GSMA and ETSI material explains why synchronized endpoint behavior, repeated registration, congestion controls, backoff, and priority handling matter in mobile cores. Those standards and guidelines define control classes; they do not establish which private timer, threshold, message, or option Docomo used during the incident. [12]-[15][18]-[20]
- The central accountability question is whether migration and rollback authority was bound to a current endpoint-population record, a representative load model, separate control-plane capacity, service-specific throttles, staged batches, abort thresholds, independent reachability probes, and retained recovery evidence.
- The Heng.lu surface is telecom continuity. Running-code behavior controls reality: a procedure can authorize a safe switchback on paper, but it cannot make an overloaded signalling plane register devices or complete calls. Accurate records matter because they make operating state testable; they do not make the declared state true.
A switchback is a new operating event
Rollback is often described as a return to safety. The idea is intuitive. If a new system behaves incorrectly, restore the old system and recover the previous condition. That description can be accurate for a local, static entity. It becomes incomplete when a distributed network and a large endpoint population have changed state while the migration was under way.
NTT Docomo's October 2021 incident exposes the difference. The operator was moving an IoT subscriber and location-information server from old equipment to new equipment. A software specification issue affected some overseas-roaming IoT behavior. The operator then switched the population back. According to Docomo's report, the procedure returned a large number of IoT devices at once, producing a mass of location-registration signals. [2][3][5]
The old server may have been familiar. The load arriving at it was not necessarily the load that existed before migration. The returning population had to generate registration work, and retry behavior could amplify load after rejection or delay. Control-plane systems and signalling switches had to process a synchronized population transition rather than a normal background flow.
That is why "the previous equipment was restored" is not a sufficient recovery statement. A distributed rollback has at least four states:
- The configuration or server placement to which the operator intends to return.
- The endpoint population that must reattach, re-register, or retry.
- The control-plane resources that must process that return.
- The voice, data, emergency, and downstream services that must become usable again.
Each state can recover on a different clock. An old server can be active while registration queues grow. A signalling switch can accept some requests while rejecting others. A restriction can be lifted while customer devices remain in an impaired state. One radio generation can recover while another remains congested. The October event included precisely that kind of staged outcome. [1][3][7]
The accountable unit is therefore not the rollback command alone. It is the whole transition from one operating population to another. Before execution, an operator should know how many endpoints may move, how quickly they can return, which signalling resources they share, which traffic can be isolated, how retries are bounded, and what evidence will stop the next batch. During recovery, the operator should observe successful registration and service use, not merely process completion.
This is not a generic lesson about being careful with changes. The network mechanism is the thesis. Remove the location-server migration, endpoint return, registration signalling, shared switches, overload controls, and staged mobile-service recovery, and the accountability argument collapses.
The impact record contains several clocks
Large incidents are often compressed into one outage duration and one affected-user number. That can make a report easier to repeat, but it can erase the operational distinction between no service, degraded service, and staged recovery.
Docomo and the Ministry published several measurements for this event. The unable-to-use interval was reported from 17:37 to 19:57 JST on 14 October, a duration of two hours and twenty minutes. Approximately one million users were estimated to have been affected by that condition. A separate difficult-to-use interval began at 16:54 on 14 October and continued until 22:00 on 15 October, a duration of 29 hours and six minutes. For the difficult-to-use condition, the operator estimated approximately 4.6 million voice users and at least 8.3 million data users. [3][7]
Those figures should not be summed into a unique-person total. A customer could appear in more than one service estimate. Voice and data populations can overlap. The methodology for estimating inability may differ from the methodology for estimating difficulty. The figures also describe conditions, not necessarily a continuous identical experience for every user.
The distinction is more than statistical caution. It reveals the recovery structure.
At 16:54, mass IoT location registration began. From 17:37, Docomo imposed controls on 4G location registration. It later relaxed restrictions by area and described sequential recovery beginning around 19:57. Yet customer difficulty continued. The operator began adjusting IoT registration volumes later that evening. It reported 5G and 4G recovery at 05:05 on 15 October and 3G recovery at 22:00. [1][3][5][7]
An incident can therefore pass several milestones:
- The most severe inability ends.
- A broad control is relaxed.
- Registration success improves in selected areas.
- Voice and data become usable for most customers.
- One radio generation returns.
- Devices that fell back to another generation transition back.
- Downstream services confirm recovery.
- Residual congestion and customer actions cease.
If an operator publishes only the earliest favorable milestone, the communication can be technically true while operationally misleading. If it waits for every possible residual symptom, it may fail to give useful interim information. The remedy is service-specific language: what is restored, for which population, by which measurement, and what remains impaired.
The Telecommunications Carriers Association's reliability and outage-communication materials provide sector context for accurate and service-specific reporting. Later guidance cannot establish exactly how Docomo communicated during October 2021, but it helps define the evidence that future notices should preserve. [16][17]
The timeline also changes how remediation should be tested. A rollback drill should not declare success when a process exits cleanly. It should measure registration queue depth, signalling-switch utilization, acceptance and rejection rates, voice-call completion, data-session establishment, emergency-call reachability, downstream MVNO status, and recovery by radio generation. The clock should stop only when the declared service objective is met.
Mass location registration became a control-plane load
A mobile device does not become usable merely because it can hear a radio signal. The network must know enough about the device and subscriber to authenticate, locate, route, and support service. Mobility management and location registration create signalling work in the core.
The public incident accounts identify that work as the immediate pressure point. When the IoT population was switched back, a large number of devices generated location-registration signals. Congestion formed between subscriber or location servers and signalling switches. Shared processing resources were consumed, and the effect spread through the nationwide network. [3][5][7]
This is a control-plane failure mechanism. The customer-visible harm appeared as voice and data difficulty, but the initiating load was not simply user payload traffic. It was the network's effort to establish or update state for endpoints.
That distinction matters because capacity planning based only on average payload volume can miss signalling risk. An IoT device may send very little application data yet still create meaningful control-plane load when many devices connect, detach, roam, reboot, or retry together. A fleet can be quiet in steady state and disruptive during a synchronized transition.
GSMA connection-efficiency guidance describes the wider class of risk. Poorly coordinated or synchronized IoT behavior can create excessive signalling, and recovery behavior can amplify load when many devices attempt to reconnect at the same time. Randomization, backoff, bounded retries, and efficient connection management are among the tools used to reduce that pressure. [12][18][19]
ETSI and 3GPP specifications describe mobility signalling and network congestion-control mechanisms, including rejection and backoff behavior. They provide a technical vocabulary for asking how registration requests are accepted, delayed, prioritized, or rejected. [14][15]
Those materials do not prove Docomo's exact configuration. The public record does not disclose every message type, timer value, rejection cause, vendor implementation, or per-node threshold in the October incident. It would be incorrect to infer a private parameter merely because a standard permits it.
They do support a concrete accountability agenda:
- Was the number of devices expected to return recorded before the switchback?
- Did the model include roaming devices and delayed retry behavior?
- What registration rate could the subscriber servers and signalling switches sustain?
- What queue, CPU, memory, or transaction threshold would stop the batch?
- Could the network instruct the IoT population to back off without applying the same restriction to ordinary users?
- Were retries randomized, or could they become synchronized again after a common rejection period?
- Did priority and emergency traffic retain access to separate capacity?
- Were the tests performed at production-representative population and signalling scale?
The value of those questions is that each can produce evidence. A population manifest, load-test result, capacity envelope, threshold configuration, staged migration log, rejection-rate graph, and call-completion probe can be retained. A general statement that a rollback plan existed cannot answer them.
Shared signalling resources expanded the blast radius
The most consequential architecture fact in the public record is the shared-resource boundary. Docomo said IoT and ordinary mobile users shared location-registration processing resources in signalling switches. It also said it could not initially regulate only the IoT population. [3][5]
That coupling allowed an endpoint transition in one service class to impair voice and data service for a much wider population. The trigger involved IoT migration. The harm crossed into a national mobile network because the control-plane resources were shared and the available restriction was not sufficiently selective.
Shared infrastructure is not inherently negligent or defective. Pooling can improve utilization, simplify operations, and provide scale. The accountability question is whether the shared resource has isolation proportional to the consequences of overload.
Isolation can take several forms:
- Separate processing capacity for populations with different retry behavior.
- Admission control that recognizes a device or subscriber class.
- Per-class queue and rate limits.
- Reserved capacity for ordinary voice, data, emergency, and priority services.
- Failure domains that prevent one migration batch from consuming national capacity.
- Independent telemetry for each population.
- A control path that remains available while service-plane congestion grows.
The Ministry's guidance required Docomo to minimize mutual impact between IoT services and voice or other communications. Docomo's response described two particularly relevant changes: separating location-registration processing resources for IoT and ordinary terminals, and adding the ability to restrict IoT location-registration signals independently. It also described network-control procedures based on observing resource utilization and adjusting restrictions. [5][8]
These are stronger remedies than an instruction to avoid future mistakes. They change who competes for capacity and who can be throttled. They move the control from a generic nationwide restriction toward a population-specific containment mechanism.
The public documents nevertheless leave proof questions open. A planned separation is not the same as deployed separation. A feature that can restrict IoT traffic is not the same as a tested threshold and operator procedure. A resource partition can be too small, share another dependency, or become stale as the device population grows.
Durable evidence would include the date and scope of deployment, the classes recognized by the control, capacity reserved for each class, load-test results, alert thresholds, exercise records, and change history. It would also show whether emergency calls and other critical services have independent operating paths or only logical priority inside the same exhausted subsystem.
This is where the Heng.lu doctrine's running-code principle is useful. A design document can record an intended boundary. The actual queues, resource consumption, rejection behavior, and completed services reveal whether the boundary exists when stressed. The record is necessary because it lets operators and reviewers compare intention with reality. It is not sovereign over the running system.
Reversal required an endpoint-population ledger
Large migrations often keep careful records of servers, software versions, interfaces, and maintenance tasks. The Docomo incident suggests that an equally important entity is the endpoint population affected by the transition.
The operator and contractor needed to know not only which subscriber or location server would be active, but which devices would be directed to it, what state they would hold, how many would return at once, and how they would behave after rejection or delay.
An accountable endpoint-population ledger would not need to identify individual customers in a public report. Internally, it should bind the migration to measurable classes:
| Population attribute | Why it matters |
|---|---|
| Device or service class | Different firmware and applications can reconnect differently |
| Domestic or roaming state | Roaming behavior can expose specification and test gaps |
| Expected active count | Defines the ordinary registration baseline |
| Maximum simultaneous return | Defines the switchback surge |
| Retry and backoff behavior | Determines whether load decays or synchronizes |
| Priority class | Protects emergency and essential service |
| Assigned server and signalling path | Reveals shared dependencies |
| Batch and cutover window | Enables bounded execution |
| Observed registration success | Shows whether the batch is healthy |
| Abort and release threshold | Prevents the next batch from advancing |
This is an operational recordkeeping function. The ledger does not own the devices or create authority merely by listing them. Its purpose is uniqueness, accuracy, transfer recording, security metadata, and continuity. A migration controller can use it to decide which population moves, prove that the expected population moved, and detect when an unplanned population returns.
Without that record, a switchback can be treated as a server operation even though its real load is generated by millions of clients. The control system sees the box being restored but not the population storm it is authorizing.
The public reports indicate that old and new equipment behavior was not fully aligned for some overseas-roaming IoT use and that Docomo and its contractor did not share the same understanding of the switchback procedure. [5][7][8] That combination points to two linked records: a specification-difference ledger and a population-transition ledger.
The first should identify every old behavior that the new software must preserve or intentionally change. The second should identify which endpoints depend on each behavior and how they move during cutover and reversal. Testing one without the other can miss the actual load and compatibility boundary.
Docomo's announced response included comparing old and new specifications and adding overseas-roaming tests. It also included clearer cutover and switchback procedures and confirmation by responsible managers. [4][5] Those controls become auditable when the comparison, test inputs, expected results, approvals, and exact procedure version are retained together.
Contractor coordination was a technical control
Outsourcing does not remove an operator's responsibility for the network it controls. It does create an interface where assumptions, procedures, and authority can diverge.
The Ministry and Docomo records describe a difference in understanding between the operator and a contractor over the switchback procedure. [5][7][8] That is not merely a communications issue. In a mobile-core migration, procedure determines which population moves, in what order, under which conditions, and who can stop or reverse the work.
The accountability model should separate actors by practical control.
NTT Docomo controlled the public mobile service, migration authorization, network architecture, shared-resource design, traffic restrictions, customer communication, and recovery declaration. It therefore carries the central duty to establish safe procedures, verify the contractor's plan, bound the population, monitor the network, and protect ordinary and critical service.
The contractor may have controlled implementation details, equipment behavior, procedure drafting, test execution, or operational steps. The public record does not disclose the complete contract or authority map. Responsibility for a specific mistake cannot be assigned beyond the published findings. The operator still needs evidence that delegated work meets its controls.
Equipment and software suppliers may control product behavior, defects, documentation, and fixes. The public material does not identify a vendor fault finding or disclose enough detail to allocate causation to a particular supplier.
IoT service providers and device makers can influence connection efficiency, retry logic, and fleet behavior. They do not control Docomo's shared signalling-switch architecture or national restriction authority. Their obligations should follow the behavior they can change.
Customers can reboot devices, follow service guidance, or design application continuity. They cannot create selective registration controls inside Docomo's core or qualify the operator's migration procedure.
The regulator can set obligations, investigate, require remediation, and promote sector learning. It does not execute the operator's cutover or run the signalling plane.
A strong operator-contractor interface turns those boundaries into a control artifact. It states who owns the endpoint inventory, who validates old and new specifications, who authorizes each batch, who watches which signal, who can halt the work, who executes rollback, and who declares service recovery. Each role should have a named alternate and a time-stamped action record.
Mutual managerial confirmation can reduce misunderstanding, but signatures alone are weak evidence. The confirmation should bind to the exact procedure, source and target software, endpoint population, predicted signalling load, thresholds, and recovery plan. Otherwise, two managers can approve the same ambiguous document.
Rollback deadlines need operating thresholds
Docomo's response described changes to rollback decision rules. Work would have a final decision time that reflected investigation and switchback duration. Material customer reports could trigger immediate reversal. Expected alarms and traffic changes would be identified in advance. [5]
These are important because delay during a high-impact change can expand the affected population. A maintenance window can create pressure to keep investigating rather than reverse. A deadline assigns value to remaining recovery time and makes indecision visible.
Time alone is not enough. A safe decision framework combines a clock with operating thresholds:
- Maximum failed or delayed registration rate.
- Maximum signalling-switch utilization.
- Maximum queue growth.
- Maximum voice-call setup failure.
- Maximum data-session establishment failure.
- Maximum emergency-call impairment.
- Maximum number of geographic areas in restriction.
- Maximum divergence between expected and observed endpoint counts.
- Maximum duration without a reliable cause classification.
- Minimum time required to reverse safely before the maintenance window ends.
Each threshold should specify its source, sampling interval, owner, and action. "High traffic" is not a trigger. "Sustained registration processing above the tested envelope for five minutes, with call completion below the service objective, stops the next batch and starts controlled reversal" is a trigger that can be audited.
The reversal itself must be bounded. If the entire population is returned simultaneously, rollback may reproduce or worsen the overload. A safer system can pause new moves, isolate the affected cohort, restore a limited batch, observe resource state, and advance only after acceptance criteria pass.
This creates a two-sided rollback plan:
- Restore the intended server or software state.
- Control the population and signalling state created by that restoration.
The first is configuration recovery. The second is service recovery. The October incident shows why both must be designed before the change begins.
Recovery must be measured at the service boundary
Operators need internal milestones. A server can be healthy. A signalling switch can return below a resource threshold. A restriction can be lifted. Those events help responders coordinate, but customers experience completed service.
For this incident, useful service-boundary evidence would include:
- Successful mobile registration by geography and radio generation.
- Voice call setup and completion.
- Data session establishment and packet delivery.
- Emergency-call completion.
- SMS or messaging delivery where relevant.
- MVNO and downstream-provider success.
- IoT fleet reconnection without renewed signalling spikes.
- Devices returning from 3G fallback to 4G or 5G.
Docomo's staged timeline shows why this matters. The acute unable-to-use period ended before the longer difficult-to-use period. 5G and 4G recovered before 3G. Some users needed device-side actions or gradual transition. [1][3][7]
An honest recovery notice should map an internal action to an external measurement. For example: registration restrictions were relaxed in specified areas; successful registrations remained above a defined rate; voice call completion recovered; data sessions were usable; one radio generation remained impaired. This avoids treating a control action as proof of its result.
IIJ's notice provides a useful second plane. IIJ reported effects and recovery for services using Docomo's network, including voice, data, M2M, and IoT. [9] A downstream operator does not see every Docomo internal state. It can show whether service dependencies outside the primary operator were usable.
The strongest recovery record would reconcile:
- Docomo's internal resource and registration measurements.
- Retail customer-service probes.
- Emergency-service evidence.
- MVNO and enterprise-IoT reports.
- Geographic and radio-generation status.
- Residual customer actions.
No one measurement is complete. Together, they can prevent a premature "restored" declaration.
Emergency calls changed the public-interest threshold
Mobile networks support ordinary private communication, but they also carry emergency calls and enable payments, logistics, transport, and asset management. Japan's Ministry emphasized those wider dependencies when it issued administrative guidance. [6]-[8]
The regulator's treatment of the incident as a serious accident matters because it moves the event beyond a private quality dispute. It establishes that the scale, duration, or service effects crossed a formal telecommunications threshold and required a documented response.
That classification should not be stretched into claims the record does not support. It does not by itself establish negligence, intent, individual fault, a damages amount, or a violation beyond the regulator's stated findings. The article does not infer those conclusions.
It does support a higher evidence standard for critical-service continuity. If emergency calls can be affected by shared control-plane congestion, the operator should be able to show:
- Which emergency-call paths depend on the affected registration resources.
- Whether priority treatment survives the overload class.
- Whether alternate networks or fixed-line routes are genuinely independent.
- How emergency organizations receive timely, specific notice.
- Which customer guidance is safe and practical during impairment.
- How exercises test the failure of ordinary and fallback paths together.
Fallback advice can be dangerous if it assumes independence that does not exist. A customer may be told to use another device, radio generation, or network, but the alternate path may share location-registration resources, backhaul, power, or an overloaded interface. The dependency map must show where the separation is physical, logical, procedural, or merely assumed.
Docomo's remediation response included communication improvements and industry sharing. The TCA materials provide a mechanism for sector guidance. [5][16][17] The durable proof is whether later exercises and notices can identify affected services rapidly, state what remains impaired, and give alternatives whose independence has been tested.
IoT is not outside the public network
The incident also challenges a common mental boundary. IoT connectivity can be treated as a specialized service separate from ordinary mobile users. Operationally, it may share subscriber systems, signalling switches, radio access, transport, identity, and control procedures with the public network.
The October outage began with an IoT server migration and affected ordinary voice and data service because that shared infrastructure mattered. [3][5][7] The IoT population was not an external workload merely consuming spare capacity. It was part of the core's control-plane state.
This has two implications.
First, IoT scale should be assessed in signalling terms, not only data volume. A meter, tracker, terminal, or embedded device may send small payloads while producing significant registration work during fleet-wide reconnection. The critical number is not only bytes per month. It is simultaneous attachments, registration attempts, retry distribution, roaming behavior, and recovery synchronization.
Second, IoT contracts and onboarding should include network-continuity controls. An operator should understand how a fleet behaves after loss of coverage, server migration, rejection, reboot, or time synchronization. Device makers and service providers should implement efficient, bounded connection behavior. Operators should protect shared resources even when devices behave badly.
GSMA guidance addresses connection efficiency and operator protection mechanisms. [12][13][18]-[20] That material supports a shared-control model:
- Device and application designers should avoid synchronized, unbounded retry.
- IoT service providers should maintain current fleet and firmware records.
- Mobile operators should identify populations, enforce admission controls, and isolate core resources.
- Roaming partners should test behavior across relevant environments.
- Critical users should understand continuity dependencies.
The duties are complementary. A device-side backoff does not excuse a shared core with no population-specific control. A selective network throttle does not excuse a fleet that ignores connection-efficiency requirements. Accountability follows each actor's practical control.
Standards define possibilities, not incident facts
Technical standards can strengthen an investigation by showing what protocol behavior and control mechanisms exist. They can also become a source of false precision when a writer infers private implementation from a general specification.
ETSI and 3GPP documents describe Evolved Packet System architecture and Non-Access Stratum behavior, including mobility management, registration-related signalling, congestion, rejection, and backoff concepts. [14][15] GSMA documents discuss connection efficiency, device behavior, network protection, filtering, priority, and abnormal signalling. [12][13][18]-[20]
From those sources, it is reasonable to ask whether registration was throttled, whether retry was randomized, whether priority classes were protected, and whether the network could isolate the IoT cohort. It is not reasonable to state that a specific timer or rejection cause was configured unless Docomo's evidence says so.
This distinction protects the article from two errors.
The first is technical invention. A plausible protocol explanation can sound authoritative while being wrong for the actual network. Private vendor implementations, software releases, roaming arrangements, and policy can alter behavior.
The second is control theater. An operator can cite standards compliance without showing that the relevant option was configured, tested, monitored, and effective under production load. Conformance to a protocol does not prove sufficient capacity or safe migration procedure.
The evidence chain should therefore move through four levels:
- The standard identifies a possible or required mechanism.
- The operator records the selected implementation and configuration.
- A representative test exercises the mechanism under the expected population and load.
- Production observations show the mechanism contained or recovered the incident class.
Only the fourth level proves running behavior. The earlier levels make that proof interpretable.
Announced remediation needs independent operating proof
Docomo's December response described a substantial program. It included comparing old and new specifications, testing overseas-roaming behavior, clarifying procedures with contractors, setting rollback decision times, defining expected alarms and traffic, adding IoT-specific regulation, separating resources, building network-control procedures, conducting exercises, improving customer communication, and sharing lessons through the industry. [4][5]
Those controls align with the failure mechanism. They address compatibility, population transition, authority, timing, isolation, overload, observability, and communication rather than relying only on training.
The open question is durability. Public reports generally describe intent and completion plans. They do not expose all production configurations or continuing test results. A control can be deployed once and later weakened by growth, software replacement, organizational change, or a new contractor.
For each announced control, the operator should retain an evidence pair:
| Announced control | Durable operating proof |
|---|---|
| Old-new specification comparison | Versioned matrix, unresolved differences, approval, and tests tied to deployed software |
| Overseas-roaming test | Representative partner and device matrix with expected and actual results |
| Shared switchback procedure | Exact procedure hash, role map, approvals, exercise, and execution log |
| Final rollback decision time | Time-stamped decision record and proof that reversal can finish within the remaining window |
| Expected alarm and traffic profile | Baseline, thresholds, alert route, response, and false-negative review |
| IoT-only registration restriction | Configuration, cohort recognition, trigger, enforcement result, and priority-service check |
| Resource separation | Architecture and load evidence showing ordinary service remains usable under IoT surge |
| Network-control exercise | Scenario, injected load, decisions, service probes, result, and remediation |
| Customer communication rule | Publication timeline, service specificity, approval path, and downstream distribution |
| Industry sharing | Guidance, entities, exercise or adoption evidence, and later revision |
This does not require publishing sensitive network configuration. Aggregated evidence can show control coverage and result while protecting exploitable details. What matters is that the operator, regulator, and qualified reviewers can distinguish a declared repair from a working repair.
A responsibility map follows practical control
Accountability becomes vague when every entity is described as jointly responsible. It becomes unfair when all consequences are assigned to the most visible brand without examining actual control. A better map links each actor to prevention, containment, evidence, communication, and recovery.
| Actor | Practical control | Evidence owed | Boundary |
|---|---|---|---|
| NTT Docomo | Migration authorization, core architecture, shared signalling capacity, restrictions, monitoring, recovery, customer notice | Exact change and procedure, population model, load envelope, thresholds, service probes, remediation proof | Cannot guarantee every device or downstream application behavior |
| Contractor | Implemented procedure, technical inputs, execution steps within delegated scope | Versioned procedure, assumptions, test results, operator confirmations, execution log | Public record does not disclose complete contractual authority |
| Equipment/software supplier | Product behavior, specifications, defect information, fixes | Release behavior, compatibility matrix, relevant defect and test evidence | No public finding here establishes vendor fault |
| IoT service/device operator | Fleet inventory, firmware, retry and connection behavior | Device-class records, connection-efficiency tests, controlled update and retry policy | Does not control Docomo's core isolation |
| MVNO/downstream provider | Customer communication, service probes, continuity planning | Timestamped impact and recovery evidence, dependency map | Does not operate Docomo's signalling switches |
| Customer or public agency | Local continuity choices and response to accurate guidance | Tested local fallback where proportionate | Cannot regulate a national core-network population |
| Regulator | Rules, investigation, remediation orders, sector learning | Findings, required controls, follow-up and proportional disclosure | Does not execute production network changes |
The table avoids transferring responsibility across control boundaries. Docomo cannot make every IoT device efficient, but it can decide whether one cohort can exhaust resources shared with ordinary voice and data. A device maker cannot isolate Docomo's signalling switches, but it can avoid synchronized unbounded retries. A regulator cannot operate the network, but it can require evidence that controls were implemented and exercised.
This is a stricter standard than blame by outcome. It asks what each actor could prevent, detect, limit, communicate, or repair, and what record demonstrates that work.
A control package for the next migration
The event can be translated into a reusable migration package. The package should be machine-checkable where possible and human-authorized where judgment is required.
1. Event and population scope
Identify the exact service, server, software, interfaces, roaming behavior, device classes, subscriber counts, geographic scope, and expected simultaneous transitions. Bind the source inventory to the approved change.
2. Specification-difference record
Compare old and new behavior. List every intentional difference and unresolved uncertainty. Connect each difference to a test and endpoint population. Do not assume that functional success for domestic devices proves roaming behavior.
3. Capacity envelope
Record sustainable and burst registration rates for subscriber servers, signalling switches, and dependent systems. Include queue and resource limits. Model normal cutover, partial failure, full reversal, synchronized retry, and delayed return.
4. Isolation proof
Show which resources are shared and which are separate. Demonstrate that the IoT cohort can be throttled without denying ordinary and priority service. Test the common dependencies that remain after logical separation.
5. Staged execution
Move a representative but bounded batch. Observe for long enough to capture retry and roaming behavior. Advance only after registration, resource, voice, data, and downstream acceptance criteria pass.
6. Halt and rollback authority
Define who can stop the work, which thresholds act automatically, the last safe decision time, and how the endpoint population will return without a surge. Preserve the decision and exact action.
7. Independent service probes
Measure completed service from more than the changed system. Include ordinary users, IoT, roaming, MVNO, emergency, and radio-generation paths where applicable.
8. Communication
Prepare service-specific notices and downstream distribution. Distinguish unable, difficult, recovering, and restored states. State which alternatives are independently tested.
9. Recovery reconciliation
Align server state, signalling capacity, registration acceptance, call completion, data sessions, radio generations, geographic status, and downstream reports. Do not declare completion from one favorable metric.
10. Post-change evidence
Retain the exact deployed bytes or procedure version, approvals, telemetry, anomalies, decisions, rollback actions, and acceptance results. Schedule a later review so controls remain current as the fleet grows.
The package is not a guarantee. It creates a falsifiable record. If an assumption fails, reviewers can identify which population, capacity, boundary, or decision was wrong and improve the next execution.
An evidence table for regulator and operator review
The following table distinguishes a document from an observed result. It does not claim that Docomo lacks each item. It identifies what would demonstrate effective control.
| Control | Retained record | Observed result | Public limit |
|---|---|---|---|
| Population inventory | Device classes, roaming state, batch membership, expected counts | Observed transitions matched the authorized cohort | Customer-level data need not be public |
| Specification comparison | Old-new behavior matrix and unresolved differences | Representative domestic and roaming tests passed | Public reports summarize rather than disclose full software detail |
| Registration capacity | Sustainable and burst envelope per resource | Peak load stayed within tested limits | Per-node graphs are not public |
| Selective admission | IoT cohort policy and trigger | IoT load was restricted without denying ordinary service | Exact policy and thresholds are private |
| Resource isolation | Architecture and shared-dependency map | Ordinary and priority service stayed usable during surge | Logical separation may retain common dependencies |
| Staged cutover | Batch plan, hold points, approvals | Each stage met service and resource criteria before advance | Public record does not show every later exercise |
| Rollback deadline | Last safe decision time and authority | Decision occurred early enough for bounded recovery | Judgment quality still requires review |
| Switchback execution | Exact sequence and endpoint-return controls | Return did not create a renewed registration storm | A clean procedure run alone is limited public evidence |
| Service probes | Voice, data, emergency, IoT, MVNO, roaming checks | Customer service met declared objectives | Samples cannot cover every user |
| Recovery declaration | Criteria and time-stamped evidence | Published status matched measured service | Residual device conditions may remain |
| Contractor interface | Role map, procedure hash, mutual confirmation | Operator and contractor executed the same understood steps | Signatures do not prove technical correctness |
| Remediation exercise | Scenario, load, decisions, results, follow-up | The 2021 failure class was contained | One exercise does not prove continuous enforcement |
The unresolved-limit column is deliberate. Accountability evidence loses value when it hides what measurement cannot prove. A load test can become stale. A sample can miss a customer class. A logical partition can share a hidden database. Naming the limit creates the next verification task.
A bounded verification agenda
The public record supports a focused set of questions.
Migration and specification
- Which old-equipment behavior for overseas-roaming IoT was absent or different in the new software?
- Which test should have exposed that difference?
- How is the current specification comparison tied to deployed versions?
- What unresolved differences can block a future migration?
Population and load
- How many devices were expected to move in each batch?
- How many returned during switchback?
- What retry and backoff behavior did the population exhibit?
- What registration rate can each dependent resource sustain?
Shared resources
- Which signalling-switch resources were shared by IoT and ordinary users?
- Which controls can now identify and restrict the IoT cohort?
- Which dependencies remain shared after resource separation?
- How is emergency and priority service protected under the same overload?
Decision authority
- What observations triggered investigation and reversal?
- What was the last safe rollback decision time?
- Did the operator and contractor use the same procedure version and role map?
- Which automatic threshold can stop the next batch without waiting for consensus?
Recovery
- When did registration success recover by geography and radio generation?
- When did voice and data service meet their objectives?
- Which downstream operators confirmed recovery?
- What residual customer actions remained after each published milestone?
Durability
- When were IoT-only throttling and resource separation deployed?
- At what production-representative scale were they tested?
- When was the same failure class last exercised?
- What evidence shows the control remains effective as the IoT population and network change?
These questions can be answered without publishing every sensitive detail. They require current, bounded evidence rather than a general assurance that lessons were learned.
Conclusion
NTT Docomo's October 2021 outage was not simply an unsuccessful IT migration. It was a mobile-network control-plane event in which a switchback caused a large endpoint population to generate location-registration signals, consumed shared signalling-switch resources, and spread congestion into ordinary voice and data service. [3][5][7]
The event demonstrates that rollback is not a return to a photograph of an earlier architecture. It is another distributed transition. The server state, endpoint state, signalling state, and customer-service state can diverge. A plan that restores the old equipment without controlling the returning population can create a new failure.
The response described by Docomo and the Ministry addresses the right surfaces: specification comparison, roaming tests, contractor procedure alignment, decision timing, population-specific restriction, resource separation, network-control exercises, and communication. [4]-[8] The remaining accountability question is whether those controls are current, deployed, representative, exercised, and effective under production load.
Running-code primacy gives the standard. The approved procedure matters, but actual registration rates, queues, resource use, throttles, completed calls, data sessions, and downstream service determine continuity. Accurate records of device populations, software behavior, assigned resources, thresholds, and recovery state make that reality testable. They do not replace it.
Responsibility should follow practical control. Docomo controlled the migration and national core. Contractors and suppliers controlled delegated implementation and product behavior within boundaries the public record does not fully disclose. IoT operators controlled fleet behavior. Downstream providers controlled their probes and notices. The regulator controlled investigation and required remediation. None of those duties cancels another.
The durable repair is an evidence chain: exact population and specification records, tested capacity, selective admission, isolated resources, staged execution, explicit halt authority, bounded switchback, independent service probes, service-specific communication, and repeated recovery exercises. That chain would turn a future rollback from an assumption of safety into a verified network operation.
Source limitations
The most detailed incident and remediation records are from NTT Docomo and Japan's Ministry of Internal Affairs and Communications. They provide authoritative operator and regulatory accounts, but they do not expose every private log, command, contract term, server model, device class, resource graph, or test result. [1]-[8]
IIJ supplies independent downstream service evidence. It cannot reconstruct every internal Docomo path or identify all affected customers. NTT Group statements acknowledge impact and group response but remain related-party evidence. [9][10]
The Docomo reliability report provides contemporaneous control context, not proof that those controls prevented or contained the October event. [11]
GSMA, ETSI, 3GPP, and TCA material defines technical and sector control classes. It does not prove that a particular timer, rejection cause, priority option, capacity threshold, communication process, or network-protection mechanism was configured by Docomo during the incident. [12]-[20]
The published estimates describe different service conditions and populations. They are not added into a count of unique people. The public record does not establish exact customer loss, every emergency-call outcome, individual fault, malicious intent, negligence, civil liability, or vendor fault. This article makes none of those claims.
Announced remediation is attributed as operator or regulator evidence. Without current independent implementation and exercise results, it is not represented as proof that every control is deployed everywhere, continuously enforced, or sufficient against the same failure class.
Sources
- https://www.docomo.ne.jp/info/network/kanto/pages/211014_00_m.html
- https://www.docomo.ne.jp/info/news_release/2021/11/10_00.html
- http://ngt.idc.nttdocomo.co.jp/20211110_10.pdf
- https://www.docomo.ne.jp/info/news_release/2021/12/28_00.html
- http://ngt.idc.nttdocomo.co.jp/20211228_00.pdf
- https://www.soumu.go.jp/menu_news/s-news/01kiban05_02000233.html
- https://www.soumu.go.jp/main_content/000779906.pdf
- https://www.soumu.go.jp/main_content/000779907.pdf
- https://www.iij.ad.jp/news/information/2021/1014.html
- https://group.ntt/en/corporate/press_conference/2021/11/211110.html
- https://www.docomo.ne.jp/english/binary/pdf/corporate/csr/about/pdf/e_csr2021w_all.pdf
- https://www.gsma.com/intelligence team/wp-content/uploads/TS.34_v7.1.pdf
- https://www.gsma.com/solutions-and-impact/industries/smart-mobility/wp-content/uploads/2017/04/CLP.14-v1.1-Network-Operators-1.pdf
- https://www.etsi.org/deliver/etsi_ts/124300_124399/124301/13.04.00_60/ts_124301v130400p.pdf
- https://www.etsi.org/deliver/etsi_ts/123400_123499/123401/16.12.00_60/ts_123401v161200p.pdf
- https://www.tca.or.jp/information/anshinkyou.html
- https://www.tca.or.jp/information/pdf/Guideline_Accident_outbreak__041.pdf
- https://www.gsma.com/solutions-and-impact/technologies/internet-of-things/gsma-iot-device-connection-efficiency-guidelines/
- https://www.gsma.com/solutions-and-impact/technologies/internet-of-things/4-iot-device-application-requirements-normative-section/
- https://www.gsma.com/solutions-and-impact/technologies/internet-of-things/annex-b-connection-efficiency-protection-mechanisms-within-mobile-networks-informative-section/
Member Briefing
Deeper Profile Context
Sign in with the right membership level to unlock the full briefing and source notes.
Only for Strategic Circle
Strategic Circle
Open to all readers. Unlock profile briefings after joining and signing in.
Join Strategic CircleOnly for Leadership Alliance
Leadership Alliance
For qualified IP-asset owners and management; sign in to unlock alliance briefings.
Join Leadership Alliance
