Summary
- The FCC found that an incorrect provisioning record omitted the appropriate Comtech IP addresses. An unrelated AT&T change later pushed that record into the live Session Border Controller whitelist, interrupting the return of 911 routing information [1].
- The affected connections had been classified as customer assets rather than infrastructure assets. The FCC said that distinction allowed the change without the more rigorous testing and off-peak controls applied to infrastructure changes [1].
- When errors crossed a threshold, a Proxy Location Routing Function reset links shared by two routing-information providers. The response that was intended to avoid one unusable path therefore interrupted the other path as well [1].
- AT&T reported that about 12,600 unique callers could not reach 911 directly during the five-hour event. The manual emergency relay center was not designed for a nationwide overflow and dropped most of the additional calls it received [1].
- AT&T later told the FCC that it reclassified the gateway links, changed alarm delivery, separated the two logical paths and added a manual VoLTE-to-3G fallback. Those are recorded remediation steps, not proof by themselves of every later operating result [1].
Keep the incident boundary exact
This article concerns the nationwide AT&T Mobility VoLTE 911 outage that began on the afternoon of March 8, 2017. It does not concern AT&T's regional 911 disruption on August 22, 2023, or the nationwide wireless outage on February 22, 2024. Those later events have different dates, mechanisms and regulatory records.
The FCC's final report is the primary factual record for the 2017 event [1]. The Bureau reviewed confidential outage reports, public comments, documents and meetings with AT&T, its routing subcontractors and public-safety entities. The report says nearly all AT&T Mobility VoLTE customers across the country lost 911 service for five hours. It records approximately 12,600 unique users who attempted to call 911 but could not reach emergency services through the traditional 911 network.
That figure needs careful attribution. It was AT&T's quantification in the regulatory record. It does not establish the number of unreported attempts, individual medical outcomes, financial losses or every local service condition. The report also notes that some calls used legacy networks or reached a backup center, and that some localities appeared unaffected. A defensible account preserves those limits.
The contemporaneous FCC public notice opened PS Docket 17-68 and described a nationwide VoLTE 911 outage [2]. The preliminary report supplied an early scope and impact assessment [3]. NENA issued a statement on the event, and APCO filed an ex parte record in the same proceeding [4][5]. These public-safety records help establish the operational context, while the final FCC report supplies the settled technical reconstruction used here.
How the routing path was supposed to work
The FCC described a multi-stage call path. A customer dialed 911 on the VoLTE network and connected through a nearby LTE cell site. AT&T's emergency-call network sent call data to one of two routing-information providers, Comtech or West. The selected provider determined the appropriate Public Safety Answering Point from the caller's geographic information, added routing metadata and returned the supplemented data to AT&T. AT&T then delivered the call through the local exchange carrier serving the relevant PSAP [1].
Two control elements mattered to the outage. A Proxy Location Routing Function, or PLRF, decided whether to send a request toward Comtech or West based on the caller's cell-sector information. Session Border Controllers, or SBCs, governed access between AT&T and the external providers. When supplemented call data returned, the SBCs checked whether it came from an approved IP address.
That approved set was a whitelist. It was a security control: incoming data should be accepted only from addresses AT&T's live 911 network was prepared to trust. AT&T also kept a provisioning-system record of those addresses. The record and the live whitelist were related, but they were not the same operational state.
This distinction is central. A provisioning record can look like an inventory or security ledger. It does not carry traffic until a change transfers it into live equipment. Accuracy therefore has to survive the transfer from recorded intent to running configuration. A correct approval process around a stale or incomplete record can still deploy the wrong result.
A dormant record error became a live routing failure
The FCC found that, before March 8, AT&T's provisioning system contained an incorrect whitelist record. It did not include the appropriate Comtech IP addresses. AT&T retained provisioning logs for 90 days but could not determine when the bad record was inserted or why. Routine inventory management had not detected the difference between the provisioning record and the live whitelist [1].
At that stage, communications with Comtech still worked because the bad record had not changed the live network. The initiating event came from an unrelated project. AT&T started a network change that pushed the incorrect record into the live SBC whitelist. Once Comtech's actual sending addresses no longer matched the approved set, AT&T rejected the returned routing data and broke the connection needed to obtain the appropriate PSAP information.
The technical chain matters because it separates four controls that are often collapsed into one. The first is data quality: did the recorded address set match the intended providers? The second is reconciliation: did anyone compare the record with the working configuration? The third is change admission: did the proposed delta receive the right review and testing? The fourth is post-change verification: did a real emergency-routing transaction still complete through both provider paths?
A hash of the submitted configuration could prove what bytes were deployed. It could not prove those bytes contained every required provider address. A successful device commit could prove that the change was accepted. It could not prove that the return traffic passed the live trust check. The relevant evidence has to connect identity, intended state, deployed state and an end-to-end service transaction.
Asset classification changed the safety case
The FCC reported that the SBC connections to the routing providers were tagged as customer assets. AT&T's infrastructure assets were subject to more rigorous failure testing and specified off-peak maintenance periods. Because these links carried the customer classification, the change could proceed without those stronger controls and during peak 911 traffic hours [1].
This was not merely a naming error. Classification selected the control regime. A link can face an external service provider and still be critical shared infrastructure if losing it affects emergency-call routing at national scale. Labels that drive review depth, maintenance timing, peer approval and negative testing are executable policy. Their accuracy deserves the same discipline as a route map or access list.
The useful test is consequence-based. What is the credible service impact if this object is wrong, missing or simultaneously reset? Which emergency or public-safety functions depend on it? Does the object sit in a shared failure domain? Can its change be exercised against a representative transaction before it reaches the live network? If the answers point to broad impact, a lower-risk label should not override that evidence.
The FCC said that more careful testing probably would have exposed the incorrect IP assignment before deployment. That is a regulatory finding about this event, not a guarantee that a particular test always catches record drift. A robust control set should include a static comparison, a negative test for an unapproved address, a positive transaction through each approved provider and an observation that the paths remain independent under error.
Redundancy was coupled by reset behavior
Loss of the Comtech return path produced errors between the SBCs and the PLRF. When those errors crossed a configured density threshold, the PLRF performed soft resets on its links to the SBCs. Comtech and West traffic used those paths, so the reset behavior interrupted access to both providers. West-supported call processing resumed when links came back, then could be interrupted again as the unresolved whitelist error generated another wave of messages [1].
The architecture had multiple routing providers, but a shared recovery action coupled them. That is a more precise diagnosis than saying there was no redundancy. Comtech and West were distinct providers with geographically diverse facilities. Yet the live control behavior allowed one provider's error condition to trigger resets on paths also needed by the other.
Redundancy should therefore be tested as a failure transition, not counted as a list of components. Two providers do not create two usable paths if one threshold, reset domain, configuration object or control node can remove both. The evidence should show that failure of provider A leaves provider B reachable, that alarms identify the affected path without suppressing the healthy path, and that automated recovery does not repeatedly widen the incident.
This is also where topology diagrams can mislead. A diagram can show separate boxes and links while omitting the control logic that acts on them together. The running relationship is defined by configuration, thresholds and state transitions. Path diversity becomes credible only when a test demonstrates isolation under the same class of fault.
The manual fallback could not absorb national load
When AT&T could not obtain the appropriate PSAP routing information, it sent calls to an Emergency Call Relay Center. Professional call takers there could ask callers for location information and manually direct them to the right PSAP. The FCC said that the center was intended for a small fraction of calls that did not route normally, not for a nationwide outage. It could not handle the additional volume and dropped the overwhelming majority of overflow calls it received [1].
Fallback capacity is part of the primary service design when the fallback is cited as a continuity control. The question is not whether a backup process exists. It is what failure set the process is sized to absorb, how quickly it activates, which data it still receives and what happens after the capacity threshold is crossed.
For a national emergency service, manual handling of every possible 911 destination would require extraordinary staffing and coordination. The FCC noted the scale of primary PSAPs as part of that constraint. The accountability requirement is not unlimited spare capacity. It is an explicit statement of the fallback boundary, a measured overload behavior and another isolation mechanism that prevents a local routing failure from immediately becoming national overflow.
The public impact record shows why that boundary matters. The FCC reported fast busy signals, repeated ringing or silence for callers. It documented examples in Orange County, Florida, while also recording jurisdictions that reported no public complaints. That mix should not be flattened into either catastrophe language or reassurance. It supports a clear operational conclusion: thousands of emergency calls did not complete through the normal path, and the fallback did not have the capacity to cover the scope.
Detection was early; diagnosis and coordination were slow
The FCC timeline records critical alarm tickets within minutes of the outage. The 911 troubleshooting team acknowledged them 16 minutes after the event began. Escalation then moved serially through the 911, VoLTE, broader service and backbone teams before the IP team was engaged. Almost five hours after the start, the IP team connected the outage time with the network change and requested rollback. Service returned three minutes later [1].
An alarm is useful only if it reaches the people who can test the relevant hypothesis. Serial escalation can be appropriate for bounded incidents, but it becomes a delay mechanism when the affected service crosses several ownership domains. The key evidence is not simply alarm generation. It is delivery to each required team, shared visibility of the change timeline, authority to compare the live state with the last change and a fast way to reverse that change.
The report also describes delayed and incomplete notification to PSAPs. Some notifications arrived hours into the event and lacked clear scope, cause or geographic detail. Public-safety agencies used social media, local media and alternative numbers to reach residents. The outage record therefore includes a communications control as well as a routing control.
Operational notification should carry at least the affected service class, known geography, start time, confidence level, current workaround and next update time. Sensitive network detail can remain protected. Ambiguity about whether legacy calls still work, whether only VoLTE is affected or which jurisdictions need alternate numbers directly reduces a PSAP's ability to respond.
What AT&T reported changing
The FCC recorded four major steps that AT&T said it had taken after the event [1]. First, AT&T reclassified the SBC connections to its 911 routing providers as infrastructure assets, bringing them under a more rigorous testing process. Second, it changed alarm delivery so the 911, VoLTE and IP teams would receive relevant errors immediately and concurrently.
Third, AT&T bifurcated the logical links between the SBCs and the PLRF. The FCC said that, had this separation existed on March 8, the Comtech problem would not have interrupted call processing supported by West. Fourth, AT&T implemented a manual process to drop VoLTE service and fall back to 3G for 911 calls during a VoLTE 911 outage.
These repairs align with the observed failure chain: stronger change control, faster cross-domain diagnosis, smaller reset domains and another service path. They should still be treated as reported remediation rather than a permanent assurance. The durable proof would include configuration identity, test results, path-isolation exercises, alarm-delivery records, fallback activation evidence and later operating history.
The strongest recurrence test would recreate the class of error without exposing callers. A test environment should omit one provider address from a candidate whitelist, confirm that the change fails admission or remains contained, verify that the other provider path stays up, deliver alarms to all relevant teams and exercise a controlled fallback. A production-safe exercise can use synthetic transactions and strict stop conditions. The result should be retained with the exact configuration and software versions tested.
The ledger is necessary, but the running network decides
The provisioning system was a recordkeeper for approved network inventory. Its record was necessary for change automation and security policy. The outage demonstrates the limit of that role. The record did not become operational truth merely because it existed in an authorized system. It became consequential when it was transferred into the live whitelist, and its accuracy could be judged only against the provider addresses and emergency transactions that had to work.
This is the Heng.lu surface of the incident. Security metadata, network identity and change history are ledgers that support continuity. They are not sovereign substitutes for the live network. Permission to deploy an object does not establish that the object matches reality. A classification label does not change the consequence of failure. A second provider on paper does not prove an isolated path.
A defensible emergency-routing record should bind five things for the same change window: the approved provider identities and addresses; the provisioning record; the candidate and deployed configuration hashes; the path-specific synthetic transaction results; and the alarms, rollback and service outcome. Any divergence should stop the change or create a named incident before public calls depend on the new state.
The same principle applies to retention. AT&T could not determine when or why the incorrect record entered the provisioning system. If the evidence needed to explain a critical change expires before the discrepancy is discovered, an operator can repair service without closing the accountability gap. Retention periods should follow the consequence and discovery window of the asset, not only routine storage cost.
A practical evidence package
Before a change, the package should identify the service owner, affected call path, provider address set, asset classification, shared dependencies, maintenance window, reviewers and expected rollback. It should include a machine-comparable difference between the working live whitelist and the proposed record. A missing approved address should be a blocking result, not a warning.
The pre-deployment test should send representative synthetic transactions through each provider path. It should verify the returned routing metadata and the ability to reach a controlled PSAP test endpoint. A fault test should remove or reject one path and prove that the other stays available without a shared reset. Fallback capacity and activation should be tested against an agreed failure set.
During deployment, operators should retain the exact configuration hash, device acknowledgements, path-specific health, alarm timestamps, transaction results and the identities of approving and observing teams. A global success status is not enough when one provider can fail while another remains healthy.
After deployment, the package should reconcile the intended and running states. It should confirm both provider paths, inspect reject counters, test end-to-end routing, validate alarm distribution and record whether any rollback threshold was crossed. The closeout should state what remains unknown rather than replacing missing evidence with confidence.
For public accountability, a carrier does not need to expose provider IP addresses or sensitive topology. It can publish the affected service surface, duration, impact range, failure class, isolation boundary, fallback performance, durable repair and recurrence-test result. That is enough to show whether the control changed without revealing material that would weaken security.
The accountability boundary
AT&T controlled the provisioning record, live whitelist, asset classification, SBC and PLRF behavior, alarm routing and carrier fallback. Comtech and West supplied routing information and operated parts of the wider emergency-service ecosystem. PSAPs controlled local public notification and alternative contact methods. The FCC assembled the cross-boundary record.
Shared operation does not erase ownership. The incorrect whitelist and shared reset behavior were within AT&T's network according to the FCC. The routing providers and PSAPs still needed timely, specific information to perform their roles. Accountability therefore includes both the initiating control and the interfaces needed for others to contain its impact.
The narrow conclusion is durable. A critical network record is trustworthy only when it is reconciled with the working configuration and a successful service transaction. Redundant providers are independent only when a realistic failure leaves one path running. A fallback is protective only within a tested capacity boundary. And a fast rollback is useful only if the teams with the right evidence can identify the responsible change in time.
Sources
- https://docs.fcc.gov/public/attachments/DOC-351492A1.pdf
- https://docs.fcc.gov/public/attachments/DA-17-277A1_Rcd.pdf
- https://docs.fcc.gov/public/attachments/DOC-344049A1.pdf
- https://www.nena.org/news/334578/NENA-Statement-on-March-8-9-1-1-Outage.htm
- https://ecfsapi.fcc.gov/file/10410294707272/APCO%20Apr2017%20ex%20parte%20-%20ATT%20Mobility%20Outages%20v2.pdf
- https://about.att.com/content/dam/snrdocs/Tips%20for%20Customers%20for%20911.pdf
Member Briefing
Deeper Profile Context
Sign in with the right membership level to unlock the full briefing and source notes.
Only for Strategic Circle
Strategic Circle
Open to all readers. Unlock profile briefings after joining and signing in.
Join Strategic CircleOnly for Leadership Alliance
Leadership Alliance
For qualified IP-asset owners and management; sign in to unlock alliance briefings.
Join Leadership Alliance
