Summary

  • TPG Telecom's planned decommissioning of a legacy 4G packet core caused an 80-minute voice outage on 15 August 2024. Of 147 emergency-call attempts originating from its network, 104 reached the emergency service by using another mobile network and 43 did not reach the required termination point.
  • The Australian regulator found that TPG knew of the significant outage at 1:22am, began a full rollback, but did not notify the emergency call person for 000 and 112 until 9:07am. That gap makes the event a practical test of change control, partial-failure monitoring and accountable handoffs.

The important failure was a chain, not a single broken link

It is tempting to describe the incident as a short mobile outage. That description is accurate but incomplete. The service disruption lasted from 12:40am to 2:00am on 15 August 2024, according to the Australian Communications and Media Authority, or ACMA. Eighty minutes is the visible window. The accountability problem extends beyond it.

TPG Telecom was carrying out a planned activity to decommission its legacy fourth-generation, or 4G, packet core. A packet core is the central part of a mobile network that decides how a phone's traffic is authenticated, routed and connected to other services. The radio towers at the edge may still be transmitting while a failure in the core prevents a device from completing a voice call. That distinction matters because a network can look partly alive even when a critical user journey is broken.

TPG told ACMA that it had conducted “mission critical testing” in preparation. The scheduled change nevertheless caused all 4G signalling links, including links carrying live traffic, to go down. A signalling link does not carry the conversation itself in the way most people imagine a phone line. It carries the instructions that allow network systems to find a subscriber, set up a call and move it to the right destination. If those instructions cannot pass, the voice service may fail before any conversation begins.

Most customers were unable to make or receive ordinary voice calls or Wi-Fi calls, including emergency calls. Some customers with particular 5G-capable handsets could still call. That mixed result made the event harder to see than a total blackout. TPG's automated processes and real-time monitoring did not immediately detect that some users could not make emergency calls over the 4G path.

The incident therefore crossed at least five control points. A planned change passed its preparation gate. Live signalling failed. Monitoring did not immediately identify the critical user impact. Engineers detected the broader outage and began a rollback. The organisation then failed to notify the external emergency-call operator promptly. Treating only the first technical trigger as “the cause” hides the rest of the system that was supposed to contain it.

What happened, minute by minute

The public record supports a clear timeline, although it does not expose every internal action.

At 12:40am, the scheduled change caused the 4G signalling links to go down. Voice and Wi-Fi calling became unavailable for most affected customers. Emergency calls were included in that loss, although some devices and call paths continued to work.

At 1:22am, TPG became aware of the significant network outage. ACMA says the issue was escalated to senior management, a full rollback began, and TPG's Emergency Management Team processes were activated. A rollback means reversing a change to return the network to a previously working state. It is not proof that every affected service has already recovered; it is the start of a controlled return.

At 2:00am, services were fully restored. That closed the availability window, but it did not close the emergency-call response. TPG conducted welfare checks on affected callers between 3:22am and 5:42am.

At 9:07am, TPG contacted the emergency call person for 000 and 112. In Australia, that role is the recognised operator that receives emergency calls and transfers them to police, fire or ambulance services. Telstra performs that role for 000 and 112. TPG sent an email at 9:07am and made a phone call at 9:25am, according to ACMA's investigation report.

This sequence creates two separate clocks. The first clock measured restoration: 80 minutes from failure to full service. The second measured notification: seven hours and 45 minutes from organisational awareness at 1:22am to contact at 9:07am. ACMA's finding concerned the second clock. The regulator concluded that TPG did not notify the emergency call person as soon as possible after becoming aware of the significant outage.

That distinction is important. An engineering team can restore service quickly relative to a complex failure and still leave a serious coordination duty unmet. Restoration and notification are not substitutes for each other. They protect different parts of the emergency-service chain.

How an emergency mobile call normally moves

For a non-specialist reader, the call path can be pictured as five connected stages.

First, the handset asks a nearby radio network for access. The radio access network is the set of towers and supporting equipment that connects the phone to the mobile operator. Second, the mobile core identifies the service and sets up the call. Third, the operator routes the emergency call towards the national emergency-call interconnect. Fourth, the emergency call person receives it. Fifth, a call taker transfers it to the relevant police, fire or ambulance organisation, with information needed to handle the call.

An ordinary diagram may show backup paths around several of these stages. But a backup line on paper does not guarantee that a phone will use it under a particular partial failure. The handset, radio network, core network and interconnection systems all have to react in compatible ways. Monitoring also has to recognise when the primary path is failing even though some other traffic remains healthy.

TPG's later submission to a Senate inquiry describes mobile emergency calls passing through its core network with multiple layers of redundancy and backup paths before being handed to a Telstra interconnect. It also describes monitoring and escalation triggers. Those statements help explain the intended architecture today. They are later, self-reported descriptions. They should not be read as proof that the same controls existed in identical form, or performed as intended, during the August 2024 outage.

The historical evidence is the observed result. All 4G signalling links went down during the change. Some users retained service. Some emergency calls moved to other networks. Other calls did not reach the required termination point. Monitoring did not immediately detect the full emergency-call impact. That running behaviour is more informative than an abstract claim that the design was redundant.

What “camp-on” did — and did not do

During the outage, 147 end users attempted emergency calls originating from the TPG network. None of those attempts succeeded through the affected TPG path. However, 104 calls “camped on” to Optus or Telstra and reached the emergency call service.

Camp-on is the emergency behaviour that allows a mobile phone to use another available carrier's radio network when its home network cannot provide the call. The caller does not need to have a normal subscription with that other carrier for this emergency use. In plain terms, a TPG customer's phone may search for another reachable network and use it solely to place the emergency call.

This mechanism reduced the number of failed calls dramatically. It is an important resilience feature. It also has boundaries.

ACMA found that only certain components of TPG's core network partially lost functionality. The radio access network did not “wilt”, the regulator's term for withdrawing in a way that prompts all devices to seek another viable carrier. Because the radio layer still appeared available, not every device moved successfully to a rival network.

Forty-three emergency calls were unsuccessful and were not carried to the relevant termination point. One caller was an international roamer. TPG conducted welfare checks on the other 42 users between 3:22am and 5:42am. Eighteen said they had not required emergency assistance. Twenty-four were referred to the relevant state law-enforcement agency for follow-up, and all had made calls after the outage. ACMA reported that agencies confirmed at least two of those users had not been experiencing an emergency.

Those facts should be read carefully. They do not establish that 43 people suffered harm. They do establish 43 failures to deliver an emergency call to the required point. They also show why an aggregate success percentage is not a complete safety measure. A fallback that works for 104 callers but not 43 is valuable and insufficient at the same time.

Camp-on should therefore be treated as a separate safety layer, not as permission to weaken the home network's controls. Its operation depends on device behaviour, radio conditions, the type of failure and the availability of another carrier. It cannot be assumed to work identically for every handset, location or partial outage.

Why partial failures are especially dangerous

A total outage is often easier to recognise. Dashboards turn red. Call volumes collapse. Alarms arrive from many systems. A partial outage may leave enough activity to reassure operators while a specific critical path has failed.

That appears to be the central monitoring lesson here. Some TPG users could still make and receive calls. Particular 5G-capable handsets retained service. Parts of the core continued functioning. The radio network did not withdraw. A system-wide availability average could therefore look better than the experience of a 4G user trying to place an emergency call.

Critical-service monitoring should ask a narrower question: can representative users complete the exact journey that matters? For emergency calling, that means more than checking whether a tower responds or whether a core process is running. It means testing whether calls from the affected access technology can be set up, routed to the emergency interconnect and acknowledged at the termination point.

The difference is similar to a building with working lights but a blocked fire exit. Measuring the electricity supply does not answer whether occupants can leave safely. In a mobile network, component health does not automatically prove end-to-end call delivery.

This is why planned changes deserve both infrastructure telemetry and user-journey checks. Infrastructure telemetry reports the state of links, processes and capacity. A user-journey check actively follows a controlled test call or an equivalent verified signal through the whole route. The two views should be reconciled before, during and after a change.

The ACMA report does not publish TPG's full alarm design or test matrix. It would be speculation to identify a particular missing alarm. The defensible conclusion is narrower: the automated processes and real-time monitoring did not immediately detect the inability of some users to make emergency calls on the 4G network. Any lesson should begin with that observed gap rather than inventing an internal cause.

Testing a change is not the same as proving the live boundary

TPG's statement that it performed mission-critical testing raises an important question: what exactly did the test prove?

A test can pass and a change can still fail for several ordinary reasons. The test environment may not contain every dependency. A procedure may validate the target system but not links that still carry live traffic. The sequence used in production may differ from the rehearsal. A partial failure may combine states that the test did not create. Operators may also receive a technically correct result that is too broad to reveal an access-specific emergency-call problem.

None of those possibilities is established as TPG's cause. They are examples of why the label “tested” should not end an accountability inquiry. The relevant evidence is the relationship among the change plan, the live boundary and the success criteria.

For a legacy-core decommissioning, the live boundary includes every signalling link that still carries traffic, even if the project considers the platform obsolete. A system does not stop being operational because a retirement plan says it should be unused. If running traffic still depends on it, it remains part of the production network.

That principle is simple enough for a nonengineer: the real network is the network that is carrying calls now. Inventories, diagrams and project milestones are records of intent. They become trustworthy only when they match observed traffic and service outcomes.

A strong decommissioning gate would therefore answer concrete questions. Which links are truly idle? How was that measured? What traffic classes were checked? Which emergency-call probes cover 4G, 5G and Wi-Fi calling? What specific signal stops the change? Who has authority to order a rollback? How is completion verified after the rollback begins? Who must be notified even while engineers are still diagnosing the fault?

The public report does not tell us how TPG answered each question. It does show that the planned activity took down live signalling. That is enough to make live-boundary evidence a central accountability issue.

Rollback was necessary, but it was only one workstream

At 1:22am, TPG escalated the outage, began a full rollback and activated its Emergency Management Team. Those actions matter. They show that the organisation moved from routine change execution into incident response.

But incident response should split into parallel workstreams. One team restores the network. Another establishes the user impact. Another handles required external communications. A fourth preserves evidence and records decisions. The same person may perform more than one role in a small incident, but the duties remain distinct.

If all attention moves to restoration, notification can wait until the network is stable. That instinct is understandable and risky. An emergency call person needs to know about a disruption while it is happening or as soon as possible after the carrier becomes aware. The information can support coordination, call handling, welfare processes and situational awareness. A notice delivered after restoration cannot perform the same function during the outage.

ACMA says TPG's own process required personnel to notify the emergency call person promptly after becoming aware of an outage or disruption. TPG later said the delay resulted from a failure to adhere to policy. That explanation points to an execution gap, not an absence of written rules.

Written policy is valuable, but it is a weak control if it depends on one person remembering a separate task during a high-pressure rollback. A stronger design connects the duty to the incident state. When an incident is classified as affecting emergency calls, the response system should open the notification task automatically, identify the accountable owner, record acknowledgement and escalate if it remains incomplete.

Automation does not remove human judgement. It makes the handoff visible. Someone can still decide what facts are safe to share, but the organisation can see whether the notice has been sent, by whom, through which channel and at what time.

The legal finding was about prompt coordination

The Telecommunications (Emergency Call Service) Determination 2019 imposed duties on carriers, carriage service providers and emergency call persons. Its stated objects included maintaining high levels of access, integrity and service continuity, ensuring controlled networks can carry emergency calls, and coordinating communications during a disruption.

Section 27 applied when a significant network outage adversely affected a controlled network or facility used to carry emergency calls or supply emergency telephone services. Paragraph 27(2)(a) required the carrier or service provider, as soon as possible after becoming aware of the outage, to notify or arrange to notify the emergency call persons for 000 and 112 and for 106. ACMA limited this investigation to TPG's notification to the emergency call person for 000 and 112.

The regulator concluded that TPG was aware of the significant outage at 1:22am. TPG contacted the relevant emergency call person at 9:07am, a significant time after service had been restored at 2:00am. ACMA found one failure to comply with paragraph 27(2)(a). It also found a related contravention of the carrier licence condition that required compliance with telecommunications law.

ACMA issued a formal warning. A warning is not the same as a court finding or a financial penalty. The article should neither minimise nor exaggerate it. The value of the enforcement record is that it defines the duty and the observed failure with precision.

The later Telecommunications (Customer Communications for Outages) Industry Standard 2024 should not be applied backwards to this event. It took effect after the August incident. It concerns communication with customers and stakeholders during qualifying outages and arose in a wider reform programme. The legal basis of ACMA's TPG finding was the emergency-call determination and the linked statutory obligations in force at the time.

Keeping those instruments separate matters. It prevents a broad narrative about “poor communications” from obscuring the exact accountable handoff: notification from a carrier that knew of a significant emergency-call outage to the recognised emergency call person.

Accountability belongs to the control system, not a guessed-at employee

The public report does not name the people who planned, approved or executed the network change. It does not identify who was supposed to notify Telstra. There is no factual basis for blaming an individual engineer or manager.

Organisational accountability can still be specific. TPG controlled the planned change, the network, the monitoring environment, the rollback process, the incident classification and the notification procedure. The question is whether those controls reliably produced safe outcomes.

This approach avoids two unhelpful extremes. The first is personal scapegoating, in which a complex control failure is reduced to a person who forgot a step. The second is vague systems language, in which nobody appears to own any decision. A useful account identifies both the institutional control surface and the decision roles without inventing names.

At minimum, the control surface included change approval, live-traffic verification, emergency-call monitoring, rollback authority, incident command and external notification. Each function should have an owner and an auditable record. If the organisation later says that policy was not followed, it should be able to show how the control changed so the same omission is less likely to recur.

TPG told ACMA that relevant personnel had been trained on emergency-call management processes and that refresher training would occur annually. Training is a reasonable corrective action. It is strongest when paired with system changes: notification checklists attached to incident severity, time-based escalation, synthetic call-path tests, duty acknowledgements and post-change verification.

The goal is not to turn every incident into a punishment exercise. It is to make responsibility legible enough that the next response does not depend on memory under pressure.

A practical change-control standard for critical call paths

The incident suggests a plain-language standard that operators, regulators and customers can understand.

First, establish the live state before changing it. A decommissioning plan should be supported by current traffic evidence, not only by an inventory that labels a platform “legacy”. If any live signalling uses the component, its removal remains a service-affecting change.

Second, define user journeys that must remain successful. For mobile voice, those journeys should cover ordinary calls, emergency calls, Wi-Fi calls and relevant access technologies. A single overall call-success rate may hide failures concentrated in 4G or a class of handsets.

Third, set stop conditions in advance. Teams should know which alarm, probe failure or user-impact signal requires a pause or rollback. The person with rollback authority should not need to negotiate that authority during the incident.

Fourth, verify the rollback. Starting a reversal is not the same as restoring every path. The team should confirm signalling, call establishment, emergency routing and monitoring recovery.

Fifth, separate restoration from notification. An emergency-call notification duty should have its own owner, timer and acknowledgement. It should proceed with the best confirmed facts available and be updated as the diagnosis improves.

Sixth, preserve a transfer record. The record should show when the organisation became aware, who classified the incident, when rollback began, when service returned, when the emergency call person was contacted and whether the contact was acknowledged. Accurate transfer records turn a verbal assurance into evidence.

Seventh, learn from partial survival. If a fallback carried 104 calls but 43 failed, the review should study both groups. What conditions allowed camp-on? What prevented it? Which monitoring signal distinguished them? This is more useful than treating fallback as either a success or a failure in the abstract.

These controls are not a claim about TPG's unpublished internal design. They are the general operational lessons supported by the public sequence.

What the record does not prove

Responsible reporting is as much about limits as findings.

The ACMA report does not say that anyone died or was injured because of the 43 unsuccessful calls. It records welfare-check outcomes and referrals, but those facts do not support a claim of harm. The event is serious because emergency calls failed to reach the required termination point, not because an unverified consequence can be added to the story.

The report does not say that all TPG customers lost service. It says most affected customers could not make or receive voice and Wi-Fi calls and that some particular 5G-capable devices continued to work.

The report does not establish that camp-on will work whenever a home network fails. It worked for 104 attempts in this event and did not work for 43. Device behaviour, radio availability and the shape of the failure matter.

The report does not publish the complete change plan, network topology, vendor configuration, testing method, alarm rules or staff roster. Any detailed reconstruction beyond the regulator's facts would be speculative.

Finally, the later TPG Senate submission is not an independent audit of the 2024 controls. It is useful because it describes the company's present explanation of routing, redundancy, monitoring and notification. Its claims should remain clearly attributed and dated.

These boundaries do not weaken the article. They make the accountability claim more durable: the known sequence alone is enough to show that live-boundary verification, partial-failure detection and emergency notification belonged to one continuity system and did not all perform as intended.

Why this matters beyond one Australian carrier

Telecommunications networks are continually changed. Old equipment is retired, software is upgraded, traffic is moved and capacity is rebalanced. Most changes succeed. The industry could not operate if every change were treated as exceptional.

The challenge is to recognise which services deserve stricter evidence. Emergency calling is one of them because failure can matter at the exact moment a user has no practical substitute. The service also crosses organisational boundaries: handset makers, radio networks, mobile cores, interconnects, the emergency call person and emergency service organisations each handle part of the path.

That means continuity cannot be reduced to ownership of a box or compliance with a checklist. It depends on accurate records of the running network, unique responsibility for each handoff, tested fallback behaviour, secure operational metadata and the ability to continue service through change.

The TPG event is useful precisely because it was not a day-long national blackout. It was a planned overnight change, a partial failure, a successful rollback and a delayed notification. Those ordinary features expose the controls that organisations are most likely to take for granted.

For customers, the lesson is not that mobile emergency calling is unreliable in general. It is that fallback mechanisms have conditions and that carriers must manage the full path. For operators, the lesson is not to avoid change. It is to prove which traffic is live, test the actual critical journey and make notification an observable incident task. For regulators, the event shows the value of publishing a precise timeline and separating a coordination breach from broader outage rhetoric.

The final accountability test is simple: after a critical network change, can the organisation show what was running, what failed, when it knew, how it restored service and when it handed the incident to every party that needed to act? If the answer depends on recollection rather than records, the continuity design is incomplete.

Sources