Summary

  • The FCC's October 2023 Notice of Apparent Liability describes two Lumen wireline 911 outages in South Dakota and North Dakota in February 2022. The notice is an allegation-and-proposal document, not a final forfeiture order.
  • In South Dakota, one signalling-path card failed and produced a loss-of-redundancy alarm. About 24 hours later, a second card failed, leaving the Pierre switch unable to complete calls that had to leave its local calling area, including affected 911 calls.
  • The South Dakota outage lasted almost five hours and potentially affected the ability of up to 14,339 Lumen wireline customers to call 911. Lumen reported within the FCC record that nobody attempted a 911 call during that outage and that no 911 calls failed.
  • In North Dakota, one of two signalling paths remained deactivated after testing. The remaining path depended on fibre transport through Chicago and Fargo. A fibre cut affected the Chicago path, while cooling problems caused equipment on the Fargo path to overheat and shut down.
  • The FCC says the North Dakota event affected 11 public safety answering points, or PSAPs—the local emergency call centres that receive 911 calls. It records 413 failed call attempts, including 49 apparent carrier tests and 364 consumer calls that did not reach a PSAP.
  • The same Chicago and Fargo transport paths also supported emergency-service trunks from Lumen's Bismarck switch into North Dakota's next-generation 911 network. Multiple originating providers had no alternative route around that ingress, or network entry point, according to the FCC record.
  • Lumen's automatic notification design initially covered centres directly served by an affected office but did not immediately identify all centres affected indirectly through the shared ingress architecture.
  • The accountability test is not whether a design document contains two lines labelled “diverse.” It is whether live configuration, alarms, trouble tickets, restoration records and notification lists identify the dependencies that actually determine whether emergency calls arrive.

What happened in South Dakota

The first outage began with a warning rather than an immediate total failure. At about 5:51 a.m. Central Standard Time on 16 February 2022, a switch card failed at Lumen's Pierre switch in South Dakota. The card provided an interface to a signalling transfer point path identified in the FCC record as Path A–St. Paul.

A signalling transfer point, or STP, is a relay or path inside Signalling System 7. SS7 is the control system that helps a traditional wireline network set up, route and complete a call. It does not carry the caller's conversation in the ordinary sense. It carries instructions that allow the network to establish and manage the connection. If the relevant signalling paths are unavailable, a telephone switch can lose the ability to complete calls that depend on those paths even when other equipment still has power.

The first card failure produced an alarm showing that Path A was no longer functioning. The alarm also told Lumen that the STP links had lost redundancy. In plain language, the switch was now relying on its remaining path. The FCC says its investigation found no record of Lumen attempting to troubleshoot the cause of the first failure at that time. That wording has a careful limit: a missing record is not proof that no person took any action outside the material the investigators reviewed. It is evidence that the public investigation did not find a documented attempt to diagnose the cause then.

About 24 hours later, at approximately 5:50 a.m. on 17 February, the switch card serving Path B–Minneapolis failed. Both STP paths were now down. According to the FCC notice, SS7 no longer functioned at the Pierre switch, so calls intended for destinations outside the local calling area could not be completed. That included 911 calls that had to leave the local area to reach next-generation 911 facilities.

Next-generation 911, or NG911, is an IP-based system designed to carry emergency calls and related information to the correct call centre. The term “next generation” does not mean the path is independent of older call-control and transport infrastructure. A call can move toward an NG911 service while still relying on wireline switches, signalling links and transport routes along the way. The South Dakota event illustrates that continuity is determined by the whole path, not the name of its newest component.

Lumen became aware of the 911 outage when the second STP path failed. The FCC record says the company dispatched a technician at 6:50 a.m., one hour after the second failure. Lumen had replacement switch cards stored onsite. The technician replaced the two failed cards and restored service at 10:43 a.m.

The outage lasted almost five hours. The FCC says it potentially affected the ability of up to 14,339 Lumen wireline customers to call 911. That is a measure of potential exposure, not a count of failed calls or harmed people. The notice also records Lumen's report that none of those customers attempted a 911 call during the outage. On that account, the South Dakota event produced no failed 911 calls.

The network failure was only one part of the incident. Lumen also had a process intended to identify affected PSAPs and notify them. The FCC notice says an automated data flow should have put enough information into a trouble ticket to identify the failed location and the relevant call centres. Instead, the ticket contained incomplete and invalid information. The automated notification distribution did not occur, and the expected manual follow-up did not correct the problem. A team member saw the notification failure but cancelled the ticket after deciding that the alert had been generated in error.

Parts of that process are redacted in the public notice. They should remain redacted in any responsible account. The record does not support guessing which internal system supplied a field, which person made each decision or what undisclosed instructions appeared on a screen.

The next day, a member of Lumen's outage-reporting team noticed that the ticket looked unusual because fields that should have contained data were blank. A later internal review concluded that the outage had affected 911 service for two South Dakota PSAPs and that notifications should have been sent. The FCC says Lumen notified those centres five days after the outage ended.

South Dakota therefore presents two dependency failures. The first was technical: one failed path remained degraded until the second path failed. The second was informational: the notification process depended on complete ticket data and correct follow-up, but those controls did not identify the affected centres when needed.

What happened in North Dakota

The North Dakota event involved more layers and two distinct outage phases. It began before any 911 call failed.

On 19 February 2022, Lumen noticed instability on one of two SS7 paths serving its Bismarck, Dickinson and Mandan switches. A technician deactivated that path for testing. The path was later restored for Dickinson but not for Bismarck and Mandan, according to the FCC notice. Those two switches continued to operate over the second path, but they no longer had a redundant STP path.

That remaining signalling path relied on two fibre transport circuits operated by a third party. One route ran through Chicago and the other through Fargo. The FCC record describes the circuits as diverse, meaning they were intended to provide separate transport options. Diversity is useful only to the extent that the routes do not fail together and do not converge on another component the service cannot avoid.

At approximately 12:18 p.m. on 21 February, a fibre cut occurred on the Chicago transport path near Henderson, Colorado. The next morning, equipment on the Fargo path began to overheat because of heating, ventilation and air-conditioning problems. The overheating caused shutdowns and affected traffic flow.

The public notice records an important visibility gap. Lumen said it was not aware at the time of either the Chicago fibre cut or the serious Fargo cooling problem. It later reported learning of the Fargo issue about 30 minutes after 911 service was restored and of the Chicago cut only after restoration. The notice's narrative and timeline differ on the precise elapsed interval for the Chicago information, so this account does not choose one figure. The agreed point is the operational one: the relevant transport faults were not known to Lumen while the emergency-call disruption was unfolding.

An operator cannot respond to a failing backup route if the state of that route is not visible in time.

At 8:15 a.m. on 22 February, Lumen received a loss-of-redundancy alarm indicating that the Chicago path was down. At 9:00 a.m., it received an alert showing an SS7 outage condition. The FCC says this loss of signalling connectivity prevented transmission of 911 calls to 11 PSAPs in western North Dakota.

This was the first phase. At about 10:45 a.m., a technician reactivated the first STP path for Bismarck and Mandan. SS7 connectivity was restored, ending that phase of the outage. It would be inaccurate to describe 10:45 a.m. as the end of the whole event.

At 11:10 a.m., the Fargo transport path shut down completely. The Chicago path was also unavailable. The failure did not again remove all SS7 communications, because the first STP path had been restored. But a different shared dependency now stopped emergency calls.

The Bismarck switch had previously served as a selective router, a component that helps direct a 911 call toward the centre responsible for the caller's location. North Dakota had moved to NG911, so that selective-routing function was no longer performed at Bismarck in the same way. Yet the ingress architecture remained. Ingress means the point where traffic enters another network.

Lumen maintained the Bismarck office as an entry point for 911 calls going into North Dakota's NG911 network. Emergency-service trunks from the Bismarck switch to that network used the same Fargo and Chicago transport paths as the second STP path. The FCC says multiple originating service providers had no alternative route around the Bismarck ingress point.

This distinction matters. Two services did not become technically identical simply because they shared transport. Signalling connectivity and emergency-service trunks perform different functions. But their continuity depended on the same two transport paths at a point the calls could not bypass. When both paths were unavailable, 911 traffic stopped at the Bismarck switch.

The third-party carrier resolved the Fargo cooling problem at approximately 4:08 p.m. Restoring the Fargo path reconnected the Bismarck switch to the NG911 network and restored 911 service. Across its two phases, the North Dakota event disrupted 911 service for more than seven hours.

The FCC says approximately 155,792 users were affected. The calls came from several originating providers and from wireline, wireless and Voice over Internet Protocol phones. Voice over Internet Protocol, or VoIP, means voice service delivered using Internet Protocol technology. This breadth reflects the Bismarck ingress dependency: the failed delivery point served traffic originating outside Lumen as well as Lumen's own customers.

In total, the notice records 413 calls to 911 that failed to complete. Forty-nine appeared to be carrier test calls. The remaining 364 were consumer calls that did not reach a PSAP. Those numbers are call attempts, not confirmed unique callers. The record does not say whether a person retried, used another service or eventually reached emergency help.

The notice also records changing descriptions of parts of the North Dakota cause sequence across Lumen's submissions. It says the company initially tied the failure to the Chicago path and later described the interaction between that cut and the Fargo cooling problem differently. The complete internal responses and supporting material are not public in the available record, and parts of the NAL are redacted. The responsible conclusion is that the public record shows a degraded signalling state, two impaired transport routes and a shared ingress dependency. It does not support inventing one simple, final internal root cause.

How a wireline 911 call reaches a local emergency centre

To a caller, the process appears to contain one step: dial 911. The network has to perform many.

The caller's telephone provider first has to recognise the call and determine where it should go. In a traditional wireline network, SS7 helps switches exchange the instructions needed to set up, route and complete that call. An STP helps relay those signalling messages. If a switch loses the signalling paths it needs, it may be unable to complete calls beyond its local area even though the telephone and parts of the switch remain available.

The call then has to move over transport links toward the emergency network. Fibre paths carry traffic between network locations. Calling two paths “diverse” generally means they are intended not to depend on the same failure domain. A failure domain is a component or condition that can take down everything inside it. Two fibres may follow different geographic routes and still share power, cooling, equipment, software, an operations process or a destination entry point.

The call must also enter the 911 system at the appropriate place. A PSAP is the local centre where trained call takers receive emergency calls and arrange the appropriate response. NG911 uses IP-based technology to carry calls and related information, but the call still needs an available ingress into that network. If every originating provider must use one entry point and that entry point has no working route, the newer system beyond it cannot receive the call.

Finally, the operator has to understand which PSAPs are affected when a component fails. A centre may not connect directly to the office where an alarm appears. It may still depend indirectly on that office because traffic from several providers passes through a shared ingress there. A notification list based only on direct service relationships can therefore miss centres whose calls are blocked elsewhere in the chain.

This is why service continuity cannot be reduced to equipment count. The question is not simply whether two cards, two signalling links or two fibre routes exist. The question is whether the complete call journey has an independent, tested alternative from the originating provider through signalling and transport to the NG911 ingress and ultimately the PSAP.

Federal rule text reflects the critical function. Section 9.4 of Title 47 of the Code of Federal Regulations requires telecommunications carriers to transmit 911 calls to a PSAP, a designated statewide default answering point or another appropriate local emergency authority. The FCC notice applies the incident-era legal framework to Lumen's conduct. The current rule text is useful context, but it should not be used to rewrite the historical rule language in effect during the 2022 events.

The same caution applies to outage notification. Current section 4.9 contains detailed reporting and notification requirements. For the two events here, the FCC notice is the controlling public account of the requirements the agency says applied and the apparent violations it alleges.

The first map: signalling, transport and the shared NG911 ingress

A conventional network diagram often shows components and links. A continuity map has to show dependencies and failure consequences.

For South Dakota, the first useful state is not “two STP paths available.” It is “Path A failed; Path B is now the only signalling route; the cause of Path A's failure has not been documented as troubleshot in the FCC investigation.” That state should connect the loss-of-redundancy alarm to the customers and emergency-call journey exposed if Path B also fails.

The second South Dakota state is “both interface cards failed.” The effect was not a vague reduction in resilience. SS7 stopped functioning at the Pierre switch, and calls that needed to leave the local calling area could not complete. The map should therefore link each card and STP path to the specific call function lost when both are unavailable.

For North Dakota, the map needs more layers. One SS7 path remained deactivated after testing. The other signalling path depended on Chicago and Fargo transport. Those same transport paths also supported the emergency-service trunks connecting the Bismarck switch to the NG911 network. The Bismarck office remained the required ingress for 911 traffic from multiple providers, and the FCC says those providers had no alternative route around it.

A diagram that shows only the two SS7 paths might suggest redundancy was restored when the technician reactivated the first path at 10:45 a.m. That was true for the first signalling-outage phase. It was not enough for end-to-end 911 continuity. When the Fargo route fully shut down at 11:10 a.m., the shared transport dependency disconnected the Bismarck ingress from the NG911 network even though the restored STP path prevented another complete SS7 loss.

This is the core difference between component redundancy and service-path redundancy. A component can have a backup while the service still has no bypass around a shared dependency. End-to-end mapping asks what happens to the actual emergency call at each point, not whether each box has a second box beside it.

The map also needs live state. A route shown as available in a design cannot be counted as protection if it remains deactivated after testing. A fibre path cannot be counted as healthy if a cut has occurred and that state is not known to the operator in time. An ingress cannot be treated as resilient merely because the network beyond it is modern.

The phrase “dependency map” does not mean the public FCC notice contains an undisclosed Lumen diagram. It is an accountability method derived from the topology and sequence the notice describes. A useful operator map would be maintained from configuration and operational records, not reconstructed only after an outage.

Such a map should answer practical questions. Which signalling links serve each switch? Which transport paths carry each link? Which emergency-service trunks share those paths? Which providers depend on the same ingress? What alternative path remains if a route, office, cooling system or third-party carrier fails? Which alarms reveal the degraded state? Who owns restoration, and how is completion verified?

The FCC record supplies enough information to ask those questions. It does not supply the complete internal configuration, carrier contracts or current answers. That boundary is important. The evidence can show why a dependency map was necessary without pretending to reveal Lumen's full map.

The second map: which PSAPs were affected directly and indirectly

Network recovery and public-safety notification depend on related but different maps.

The first map follows the call. The second follows the consequence: if a component or ingress fails, which PSAPs will stop receiving calls, whether they are directly served by the affected office or depend on it indirectly?

South Dakota shows what happens when the data needed to build that consequence map does not arrive in the trouble ticket. The automated flow produced incomplete and invalid information. Lumen's systems could not identify the affected switch location and associated PSAP impacts. Automatic distribution failed, and manual follow-up did not repair the gap. Two centres were eventually notified five days after the outage.

North Dakota shows a different limitation. At 9:07 a.m., Lumen automatically notified two of the 11 affected PSAPs. Between 9:32 and 9:53, it notified three more. The remaining six were notified between 12:21 and 12:30 p.m.

The FCC says Lumen's automated design sent notifications when a PSAP was directly served by the office experiencing an SS7 outage. It did not initially capture all centres affected indirectly because the Bismarck switch was the ingress for their 911 traffic. Following the event, the notice says Lumen modified the design so automated notifications would reach additional indirectly affected centres. Parts of the description are redacted.

The sequence shows why a customer or facility list is not enough. A PSAP can lose calls because a network it does not directly contract with, or an office that does not directly serve it, sits on the path used by the originating providers in its region. The notification map has to be built from service dependency, not only from account relationships.

That map should also change when the network changes. A legacy selective router may stop performing its old routing function while the office remains an ingress. If the notification logic is updated only for the new service name and not for the surviving physical and operational dependency, the system may report an incomplete set of affected centres.

Accurate notification matters even when the operator cannot immediately restore service. A PSAP that knows calls may not arrive can coordinate public information about alternative contact methods. The public FCC record does not establish what alternatives were available to each North Dakota caller or what each PSAP did after receiving notice. It does establish that timely knowledge is part of managing the outage's effects.

The later notification-design change should be described with restraint. The FCC notice records that Lumen modified the design. The available public sources do not independently verify when the change was fully implemented, how it was tested, whether every indirect dependency was captured or how it performs today. A reported change is not the same as evidence of continuing effectiveness.

Alarms, restoration and the cost of an unexamined degraded state

The two incidents make degraded state as important as total outage.

In South Dakota, the first card failure did not immediately stop 911 calls. It removed redundancy. That may look like a smaller event on a dashboard because service continues. Operationally, it changes the risk of every subsequent failure. The remaining card is no longer one of two protections; it is the only path keeping the relevant signalling function available.

A loss-of-redundancy alarm therefore needs a defined outcome. The public record supports asking whether it produced diagnosis, a restoration deadline, escalation, a documented risk decision or verification that the remaining route was genuinely independent. The FCC says its investigation found no record of troubleshooting the first card failure at that time. It does not publish Lumen's complete alarm policy or staffing model, so a broader conclusion about every employee or procedure would exceed the evidence.

North Dakota contains several degraded states. One STP path remained deactivated after testing. A fibre cut affected the Chicago route. Cooling problems affected the Fargo route. The FCC says Lumen did not learn of the transport problems until after service was restored, despite receiving a loss-of-redundancy alarm at 8:15 a.m. and an SS7 outage alert at 9:00 a.m.

An alarm is useful only when its meaning is connected to the service at risk. “Chicago path down” is a component statement. “The active SS7 route and the NG911 ingress both depend on the remaining Fargo path” is a continuity statement. The second form tells an operator why the condition deserves urgent attention and what will happen if the final dependency fails.

Restoration records need the same discipline. At 10:45 a.m., reactivating the first STP path ended the first phase of the North Dakota outage. At 11:10 a.m., the Bismarck switch lost its connection to the NG911 network when the Fargo path shut down. A closeout process focused only on restored SS7 connectivity could have declared success while the end-to-end emergency service remained exposed.

The correct restoration unit is the user journey. For 911, that means a test or operational proof that a call can be set up, transported, admitted to the emergency network and delivered to the intended PSAP. It also means confirming that notification systems identify every directly and indirectly affected centre.

Trouble tickets form another evidence layer. A ticket should preserve a unique event identifier, affected components, time zone, service impact, known cause, restoration estimate and notification status. The current version of section 4.9 contains detailed examples of material outage information, including unique identifiers and geographic/service effects. For this historical incident, the FCC notice remains the source for what the agency says went wrong in Lumen's ticket and notification process.

The accountability lesson is not that every degraded alarm must trigger the same response. Networks generate alarms of different severity, and operators have to prioritise. The lesson is that the response should be tied to the actual dependency and recorded. If a path supports emergency-call signalling or a non-bypassable ingress, the decision to leave it degraded should be visible, owned and testable.

This is where records and running systems have different roles. An FCC notice, alarm log or trouble ticket can document what appeared to happen and who was expected to respond. None of those records keeps a call moving. Operational truth comes from the live configuration and from tests showing that the service continues through the claimed backup path. Records are evidence of that reality; they are not a substitute for it.

Who was affected—and what the figures do not prove

The South Dakota and North Dakota impact figures describe different things and should not be combined.

In South Dakota, the outage potentially affected the ability of up to 14,339 Lumen wireline customers to call 911. Lumen reported no attempted and no failed 911 calls during the outage. The figure is therefore a maximum potentially exposed customer count, not a call count.

In North Dakota, the FCC says approximately 155,792 users were affected. The ingress served traffic from multiple originating providers, including Lumen, and from wireline, wireless and VoIP phones. The notice separately records 413 failed 911 call attempts: 49 apparent carrier tests and 364 consumer calls that did not reach a PSAP.

A call attempt is not necessarily a person. One caller may try more than once. More than one person may use the same line. A carrier may generate test calls. The public record does not identify unique callers, and it does not say whether a consumer later connected through another attempt or method.

The record also does not establish that any failed call caused a death, injury, delayed dispatch, property loss or other individual outcome. That absence should not be used to minimise the continuity failure. Emergency calls did not reach their intended centres. It should instead prevent a factual account from adding a dramatic human story the evidence does not provide.

The same restraint applies to geographic scope. The FCC's account concerns two Lumen wireline-network events in South and North Dakota. It does not establish that every Lumen service failed, that every customer in either state lost connectivity or that the same condition existed across the company's national network.

What the FCC notice alleges and what it does not finally decide

The legal document at the centre of this case is FCC 23-81, adopted on 12 October 2023 and released on 17 October 2023. It is a Notice of Apparent Liability for Forfeiture, commonly shortened to NAL.

An NAL tells a company how the FCC believes it apparently violated the law and proposes a monetary forfeiture. It gives the company an opportunity to respond. It is not the same as a later forfeiture order, consent decree, court judgment or criminal conviction.

FCC 23-81 proposes a total forfeiture of $867,000. The notice says Lumen apparently willfully and repeatedly violated sections 4.9 and 9.4 of the Commission's rules. It attributes apparent notification violations to both outages and apparent 911-call-transmission violations to the North Dakota event.

The ordering clauses told Lumen to pay the proposed amount within 30 days or file a written statement seeking reduction or cancellation. The presence of payment instructions does not prove that Lumen paid. It describes the choices and procedure following the proposed action.

The FCC's accompanying announcement states explicitly that the allegations and proposed sanctions in an NAL are not final Commission actions. The announcement is helpful for explaining that distinction, but it is an unofficial, derivative summary. FCC 23-81 itself is the official detailed incident and enforcement record used here.

A later FCC fiscal-year 2024 financial report still described the Lumen action as a proposed $867,000 fine and as an NAL. A bounded search of official records for FCC 23-81, the file number and the NAL account number did not locate a later forfeiture order, consent decree, settlement, payment record, reduction order or cancellation order. A negative search is not proof that no non-indexed or non-public record exists. Because legal status can change, any current account needs a fresh exact-identifier check.

The precise description is therefore that the FCC notice alleges apparent violations and proposes an $867,000 forfeiture. When the notice relays the company's position, the attribution is that Lumen reported it. The available record does not establish that the FCC finally fined Lumen $867,000, that Lumen paid or settled the matter, that the company admitted the apparent violations or that a court found it liable.

The distinction is more than legal caution. A proposed enforcement record can document a detailed accountability case without becoming proof of final disposition. It can also describe a remediation claim without proving the control works. Keeping those states separate makes the article more useful: readers can see what the public evidence establishes, what the agency alleges and what remains to be verified.

What the public record does not show

The FCC notice is detailed, but it is still one public incident-record origin. The company's underlying responses and supporting documents are cited as on file with the FCC but are not fully reproduced. Parts of the notice are redacted. The FCC announcement is derived from the same action, and the record-copy PDF is another version of the same document. They do not create independent incident reconstructions.

The available public sources do not establish the complete identity, contractual responsibility or fault allocation of every third-party transport provider. They do not reveal every physical route, shared facility, alarm threshold or service-level agreement. Calling the paths Chicago and Fargo does not prove their complete geography or every component they did or did not share.

The exact way traffic degraded before each alarm is not fully public. Lumen's descriptions of the North Dakota cause sequence changed across submissions, and the company said it could not fully explain how the first phase occurred. The fair account preserves that uncertainty rather than selecting the cleanest version and presenting it as settled fact.

The record also does not establish why the first South Dakota failure was not documented as troubleshot, why the North Dakota STP path remained deactivated, or why each expected manual notification step did not occur. It would be speculation to infer intent, discipline, staffing, competence or a complete organisational root cause.

Lumen's current South Dakota customer story describes NG911, hosted call handling, MPLS/IP VPN connectivity and an upgrade involving 28 PSAPs. That operator page helps explain the kind of system involved. It is marketing and customer-story context, not an outage postmortem, corroboration or proof that the later notification change was effective.

The reviewed North Dakota state pages explain public 911 services but do not document this incident.

These limits do not make the case unusable. They define the line between a public accountability analysis and a fictional internal investigation. The available record supports a strong continuity question because it describes the call path, the degraded states, the shared ingress, the impact and the notification sequence. It does not support filling every blank with certainty.

What evidence would show that the dependency map works

The best response to an incomplete public record is not a broader claim. It is a more precise list of evidence that would test the continuity thesis.

First, a current dependency diagram should be tied to live configuration. It should identify each switch, STP path, transport circuit, emergency-service trunk and NG911 ingress used by the service. The operator should be able to show how the diagram is updated when a route is deactivated, a circuit changes carrier or an office changes function.

The important property is not visual polish. It is agreement with operation. A diagram that shows a path as active while configuration leaves it disabled is false at the moment it matters. Useful evidence would compare the recorded map with device state, circuit inventory and observed traffic.

Second, the operator should identify failure domains. A failure-domain review asks which supposedly separate paths share equipment, power, cooling, buildings, software, management systems, staff procedures, third-party carriers or destination ingress. Two routes can be geographically different and still depend on one non-bypassable entry point.

For the North Dakota architecture described by the FCC, such a review would have to connect the Chicago and Fargo transport routes to both the remaining SS7 path and the emergency-service trunks. It would also identify the consequence of losing the Bismarck ingress for every originating provider and PSAP that relied on it.

Third, controlled failure tests should exercise the complete service path. A test might remove one signalling route, then verify that an emergency call still completes through the other. It might simulate the loss of one transport path and confirm that both signalling and NG911 ingress remain available. It should also test the case where a path was intentionally deactivated for maintenance and not restored.

The purpose is not to reproduce a dangerous public outage. It is to create safe evidence about the designed backup. Results should record the exact configuration, test call path, time, expected outcome, actual outcome and any manual action required. A successful component ping is not enough if the emergency call still cannot reach a PSAP.

Fourth, alarm-to-action traces should connect a degraded state to an owned response. For each loss-of-redundancy alarm, the record should show when it was generated, how it was interpreted, which dependency and service were at risk, who accepted the event, what deadline applied, what troubleshooting occurred and how restoration was verified.

An alarm trace should also show exceptions. If the operator decides that a degraded path can remain in service, the decision should state why, for how long and what compensating control protects 911 continuity. That makes risk acceptance auditable without pretending every alarm requires the same escalation.

Fifth, maintenance closeout should verify restored state. If a technician deactivates a signalling path for testing, the work cannot be considered complete merely because testing ended. The closeout should confirm that every affected switch has the expected paths active, that redundancy alarms have cleared and that a representative call completes.

Sixth, third-party visibility should be tested. The FCC record says Lumen learned of the Fargo cooling issue and Chicago cut after the outage ended. Evidence of improvement would show how transport-provider faults become visible to Lumen, how circuit identifiers map to emergency-service dependencies, how status is escalated and how the operator verifies restoration rather than relying only on a carrier notice.

Seventh, trouble-ticket quality should be measured. Required fields should be machine-validated where possible so incomplete or invalid location data cannot silently remove the affected PSAP list. A failed automated data flow should create an unmistakable exception with a named owner and a timer. Cancelling a notification ticket should require evidence that the service impact was checked, not merely an assumption that incomplete data means a false alarm.

Eighth, the PSAP notification map should be generated from dependencies. For each affected office or ingress, the system should identify centres served directly and indirectly. The test should remove a shared ingress in a controlled environment and confirm that the notification fan-out includes every centre whose calls would be blocked.

Notification testing should record delivery as well as generation. The operator needs evidence that the correct contact was reached, that the message contained the available material information and that follow-up updates were sent as the incident changed. A generated message in a queue is not proof that a call centre received usable notice.

Ninth, restoration should be measured end to end. The North Dakota event demonstrates why one component's recovery can end one phase without ending the service risk. An incident record should distinguish restored signalling, restored transport, restored ingress, successful test call, restored notification capability and final service closeout.

Tenth, recurrence metrics should be narrow enough to reveal whether the same weakness returns. Useful measures could include time spent in a one-path signalling state, alarms without documented diagnosis, maintenance tasks closed with a path still disabled, ticket-data validation failures, indirect PSAPs omitted from notification fan-out, and end-to-end failover tests that do not complete as expected.

These measures are examples of evidence, not claims about Lumen's current system. The available record does not show whether the company now produces them. They illustrate the standard by which a reported change can become a demonstrated safeguard.

The broader accountability principle is simple. A regulator can preserve a record of apparent responsibility. A diagram can describe intended architecture. A trouble ticket can document a response. None of them is the service itself. For emergency-call continuity, the decisive evidence is that the running network survives the tested failure, that operators can see the degraded state, and that every affected call centre is identified when the path does not survive.

That is what a real dependency map should do. It should not merely explain an outage after the fact. It should make the next hidden convergence visible before a caller discovers it by dialling 911.

Sources