Summary
- On 1 March 2024, a surge of registration activity from medical alert devices generated mobile-location data and overloaded both the primary and secondary databases used by Telstra's Triple Zero platform. A latent software fault stopped the databases from recovering automatically.
- The disruption affected 494 calls. A backup list enabled 346 live transfers, 127 calls entered a phone-or-email escalation path that the regulator did not regard as a live transfer, and 21 callers said assistance was no longer required. The ACMA found 473 breaches, not 494.
- Eight of 24 backup emergency-service phone numbers were wrong. The incident shows that a separate database is not a working fallback merely because it exists: its destinations, call path, information handoff and operating procedure must all work under the same failed conditions.
What the Triple Zero platform does
A person who dials 000 or 112 is not normally expected to know the telephone number of the nearest police, fire or ambulance communications centre. The national emergency-call person receives the call, identifies the type and location of assistance required, and transfers the live call to the relevant emergency service organisation.
That transfer is more than a directory lookup. The call must remain connected while it moves from the national entry point to the organisation that can dispatch help. The receiving organisation also needs useful information. Depending on what the emergency-call platform has, this can include the public number that made the call, the customer's name, a service address and mobile-location information.
For a non-specialist, the service can be tested with three questions. Did the caller remain live during the handoff? Did the handoff reach the right emergency service? Did the useful caller information already available to the platform travel with the call? A system may pass one question and fail another.
Telstra is the emergency call person for 000 and 112 in Australia. That role gives it a central position in a service that crosses carrier networks, location systems, call-centre technology and state or territory emergency organisations. The public experience is a short number and a conversation. The infrastructure behind it is a chain of databases, signalling, contact records, trained call-takers and agreed fallback procedures.
The March 2024 incident matters because the chain did not fail in only one place. The main information platform became unresponsive. The backup transfer list contained incorrect destinations. A manual escalation path changed the service from a live transfer into a later callback. Some calls that were transferred did not carry all information the platform had. Each gap affected a different part of the promised outcome.
How the disruption began
The Australian Communications and Media Authority's final report says a large spike occurred when medical alert devices attached to the emergency-call signalling channel on Telstra's mobile network. The devices were reconnecting after being rebooted. They were not placing emergency calls at the time; they were registering in preparation for any emergency call they might later need to make.
Those attachments caused Push MoLI data to be generated. MoLI means mobile location information. In ordinary language, it is data that helps an emergency service understand where a mobile caller may be. Location can be vital when a caller cannot give a precise address, is moving, becomes disconnected or is unfamiliar with the area.
The surge caused both the primary and secondary databases that stored this data in the Triple Zero platform to exceed the maximum number of concurrent data sessions they could handle. A latent software fault then prevented automatic recovery and made the databases unresponsive. Telstra also told the regulator that other work, including security and mandatory-obligation tasks, may have coincided with the peak and contributed to web-server processor overload.
The careful wording is important. The public evidence supports a chain involving registration load, database-session limits and a previously undetected software fault. It does not support a claim that the medical alert devices attacked the service, were defective in every respect or made hundreds of emergency calls. They supplied an unusual workload that the platform failed to absorb safely.
The affected databases ordinarily provided calling-line information and stored the primary contact lists for emergency service organisations. During the disruption, call-takers could see a phone number on the screen, but the fuller calling-line information was unavailable. The usual transfer destinations could not be selected from the main system.
This coupling widened the failure. One platform area held both information about the caller and the primary list used to route the call. When the databases stopped responding, the call centre lost two different capabilities at once: enriched caller context and its normal destination-selection path.
The 494 calls and the 473 breaches are different counts
The disruption affected 494 calls. That number describes the calls received during the relevant period and affected by the platform condition. It is not the regulator's breach total.
Call-takers activated a backup process. They asked callers for the emergency location and used a contact list held in a separate database. That path successfully transferred 346 calls to emergency service organisations while the calls remained live.
Another 127 calls could not be transferred through the backup list and went into a manual escalation process. The call-taker recorded the caller's number, location and required organisation, passed the details to a supervisor, and the supervisor relayed the information by phone or email. The emergency organisation could then call the person back.
The ACMA concluded that this process did not meet the legal requirement to transfer the call. In the regulator's reading, a transfer meant switching the live call to the relevant emergency service organisation. Sending details so that the organisation could make a later callback was a useful incident response, but it was not the same service outcome.
The remaining 21 callers said that assistance was not required. The ACMA concluded that the transfer obligation did not apply to those calls because the relevant conditions were no longer met. This is why it would be wrong to state that all 494 affected calls were transfer breaches.
The regulator found 127 failures to transfer. It found a further 346 breaches for the calls that were transferred because Telstra did not make available the customer name and the most precise location information it had. The public number remained visible and was provided. The result was 473 findings in total: 127 plus 346.
These counts should not be turned into a simple victim number. They classify call-handling outcomes against specific duties. Public sources do not establish that every caller suffered the same delay, that every call concerned the same kind of emergency or that the regulatory count measures every human consequence.
Why eight wrong numbers defeated part of the backup
The separate backup list contained 24 emergency-service contact phone numbers. Eight were incorrect. Calls intended for those organisations could not be transferred using the planned backup path.
This is a remarkably ordinary failure inside a technically complex service. The main platform involved mobile-location signalling, databases and software recovery. The fallback problem included a contact list. Complexity at one layer does not make basic record accuracy less important. It often makes it more important because the simple record is what staff depend on when automation is unavailable.
A contact number can become wrong for several reasons. An organisation may change a line, a routing arrangement may be replaced, a record may be transcribed incorrectly or an update may not reach every copy. The public report does not identify the history of each of the eight entries, so assigning a single cause would be speculation.
The control requirement is nevertheless clear. A critical contact list needs an owner, a source of truth, a change procedure, a review schedule and a safe way to test reachability. The test cannot stop at confirming that digits are present in a table. It must show that the number reaches the intended operating desk and that staff know how to complete a live handoff through it.
During the incident, a changed escalation email address for Triple Zero Victoria was initially transcribed incorrectly. The error took 13 minutes to correct and delayed the provision of some information. This was a separate error inside the emergency workaround, not evidence that every call used the wrong address.
The sequence shows how fallbacks become fragile. A primary automated path fails. Staff move to a secondary list. Some entries do not work. Staff create a manual phone-or-email path. A newly supplied address is entered incorrectly under pressure. Each extra stage adds delay and another opportunity for information to diverge.
Good continuity design tries to reduce that improvisation before an incident. It keeps the alternate path current, rehearsed and close to the original service outcome. It also assumes that people under pressure need simple, verified choices rather than a collection of untested contacts.
A callback is helpful, but it is not a live transfer
To someone outside emergency communications, the difference may sound procedural. If an organisation receives the caller's details and calls back, has the message not reached the right place?
Sometimes it has, but the risk has changed. A live transfer preserves the existing connection. A callback depends on the caller's device remaining reachable, the network accepting a new call, the caller answering, and the relayed number being correct. The person may have moved, lost signal, become unable to speak or used a device that cannot accept the return call in the same way.
The callback process can also separate context from conversation. The supervisor relays a short record; the emergency organisation then reconstructs the situation with the caller. That may be the best available response during a failure, but it is not equivalent to joining the organisation to the live call.
The regulator's distinction is therefore operational, not merely semantic. The service promised by the transfer duty has a continuity property: the person remains on the line while responsibility moves to the emergency organisation. A later callback is a new communications attempt.
This does not mean the call-takers and supervisors did nothing useful. The evidence shows that they used a manual process when the backup list failed. Accountability should distinguish frontline mitigation from the system conditions that made mitigation necessary. Staff can respond diligently and the organisation can still have failed to provide the required service.
The lesson applies beyond emergency calls. A bank can route a fraud case to a callback queue, a cloud provider can email a customer after a control channel fails, and a hospital can record a message for a clinician. Each may reduce harm. None should be reported as equivalent to a live, verified handoff if the original service promised continuity.
Location and identity information are part of the service
The 346 calls transferred through the backup list reached emergency service organisations. Why, then, did the ACMA record a breach for each one?
At the time, section 51 of the Emergency Call Service Determination required the emergency call person to make available as much of three categories of information as it had: the most precise available location, the customer's name, and the public number from which the call was made.
During the disruption, the public number was visible because it was generated by the network and displayed to call-takers. The customer name and fuller location information held by the platform were not provided with the 346 transferred calls. ACMA treated each missing information handoff as a separate contravention.
The distinction matters because connection and context solve different problems. Reaching the police, fire or ambulance communications centre is essential. Knowing where to send assistance can be equally essential. A caller may describe a location accurately, but the platform's information can confirm, refine or preserve it if the call drops.
Location data is not infallible, and the sources do not say that it should replace what a caller says. The duty is to make available the most precise information the emergency-call person has. The receiving organisation can evaluate that information alongside the conversation and its own procedures.
Customer name, public number and location are also different fields. Reporting that the 346 calls carried “no information” would be wrong: the public number remained available. Reporting only that the calls were connected would also be incomplete: the other information available to the platform did not travel with them.
This is a useful infrastructure principle. A service is not always just a packet, call or transaction arriving at a destination. Metadata can be part of the promised outcome. Backup design must identify which metadata is safety-critical and preserve it through the alternate path.
Primary and secondary do not automatically mean independent
The ACMA report says both the primary and secondary databases exceeded their concurrent-session limits and became unresponsive after the latent fault was triggered. That result challenges a common reading of redundancy.
Two systems can exist and still share a failure mode. They may receive the same workload, run the same software, depend on the same recovery logic or compete for the same surrounding resources. A label such as “secondary” tells us about intended role. It does not prove operational independence.
Real redundancy is measured by behaviour under failure. Can the alternate component accept the load that appears when the primary is unhealthy? Does it fail differently? Can it recover without the same stuck condition? Is the handover automatic, observable and reversible? Has the organisation tested those questions with realistic peaks rather than ordinary averages?
The March event also shows why capacity and recovery must be tested together. A system can exceed a connection limit and reject new work while continuing to recover. The more dangerous state is when overload exposes a software fault that prevents recovery. The fallback then needs to survive not only the initial peak but the extended period during which the failed component remains stuck.
Telstra reported increasing database connection capacity, adding monitoring and deploying a software correction. Those are relevant responses. The public article does not provide independent test results, current capacity values or evidence for every workload shape. The proper conclusion is that Telstra reported specific remediation, not that all future risk has been eliminated.
For leaders, the practical question is not “Do we have two databases?” It is “Which failure assumptions are genuinely different, and what test proves that the full service continues when one assumption is wrong?”
The list is a ledger, not a working route
Operations teams need records. Contact lists, inventories, routing tables and service registries help people agree on what exists and where work should go. Without them, a national service would depend on memory and improvisation.
But a record is a ledger of intended state. It cannot make an outdated destination answer a call. It cannot keep the caller connected. It cannot attach location data that the fallback path does not carry. The operational authority of a contact entry comes from a reachable endpoint and a tested procedure, not from the fact that the entry is stored in a separate database.
This is the reality-layer test. Compare the record with the running service. Dial or safely probe the alternate number. Confirm which organisation answers. Exercise the transfer workflow. Verify that the expected caller information appears. Record the result and the date. If the test conflicts with the list, the live evidence wins and the list must be corrected.
The same principle applies to the primary and secondary labels. Architecture diagrams and asset records are useful, but the running code determines whether the two paths share a session limit, software fault or recovery mechanism. A resilience claim should be grounded in observed failover, not diagram geometry.
Accurate records are still essential. The lesson is not to distrust every registry. It is to keep the registry within its proper role: a controlled, auditable record that is continuously reconciled with the service it describes.
What Telstra and the regulator said happened next
The ACMA announced that Telstra paid a penalty of more than AUD 3 million. The regulator also noted that Telstra updated its backup phone list and appointed an independent consultant to review the incident.
Telstra's November 2024 account listed a broader set of actions. The company said it increased database connection capacity, introduced additional monitoring and notifications, updated work instructions, paused changes to the Triple Zero platforms during the investigation, corrected the eight contact numbers and scheduled regular reviews.
Telstra also said it reproduced the software issue in a laboratory, worked with enterprise customers responsible for the medical alert devices, tested and deployed a software change, altered the network to reduce registration behaviour not connected with an emergency call, reviewed end-to-end monitoring and updated business-continuity processes.
These statements matter because they map to the observed failure layers: load, software recovery, device-registration behaviour, monitoring, contact accuracy and operating procedure. They are not, by themselves, a public independent assurance report. A later reviewer would still want evidence that the controls operate: test records, monitoring samples, failover exercises, contact confirmations and closure of independent recommendations.
The regulator also acknowledged Telstra's history of compliance in the national role, public communication during the outage and immediate actions. Including that context prevents the analysis from becoming a one-sided corporate indictment. It does not erase the 473 findings.
Accountability is most useful when it preserves both facts. An organisation can communicate openly and remediate after an incident, while the incident still demonstrates that important preventive and fallback controls failed.
Twelve controls for an emergency-call fallback
First, define the service outcome. The fallback must say whether it preserves a live transfer, destination selection, caller number, customer name, location information and audit trail. “Calls can still be handled” is too vague.
Second, separate failure domains. Primary and secondary components should not silently share every session limit, software defect, deployment path and recovery dependency.
Third, test peak behaviour. Capacity exercises should include the workload created during failover and unusual signalling events, not only ordinary call volume.
Fourth, test recovery from overload. A component that rejects excess work but recovers predictably is different from one that becomes stuck.
Fifth, treat contact records as controlled infrastructure. Every destination needs an owner, authoritative source, change history and review date.
Sixth, prove reachability. A scheduled, authorised test should confirm that each alternate number reaches the intended emergency organisation and supports the expected transfer.
Seventh, preserve live-call continuity. If the planned fallback changes a live transfer into a callback, that limitation should be explicit, risk-assessed and improved rather than hidden by the word “escalation”.
Eighth, preserve critical metadata. The alternate path should identify how the public number, customer name and most precise available location reach the emergency organisation.
Ninth, rehearse human steps. Staff should be able to use the fallback under time pressure without searching for unverified addresses or inventing a new workflow.
Tenth, monitor the user journey. Database health is not enough. A controlled test should prove that a call can be received, classified, transferred and accompanied by the expected information.
Eleventh, keep independent communication channels current. When the technical handoff fails, supervisors need verified ways to coordinate with each emergency organisation, including out-of-band contacts.
Twelfth, close the loop after every test and incident. Correct records, assign findings, verify remediation and repeat the end-to-end exercise.
These controls do not promise zero incidents. They make it harder for a single overload, stale record or transcription error to turn a technical fault into a wider loss of emergency-call continuity.
Questions that public agencies and large customers can ask
Most customers cannot inspect the national emergency-call platform. Public agencies and large organisations can still ask useful questions of their communications providers and of their own continuity teams.
What happens to emergency calling when customer identity or location systems are unavailable? Which information remains visible? Does a backup transfer preserve the live call? How often are alternate destinations verified with the receiving organisation?
Are the primary and secondary systems independent in software, capacity and recovery, or only in name? Which recent exercise proved the distinction? Did the test include the extra load created when devices reconnect after an outage or restart?
How are changes to emergency-service contact details propagated? Is there one accountable owner? Does a second person verify transcription? Can staff see when an entry was last tested rather than merely last edited?
How is a degraded mode described to incident leaders and the public? A dashboard should distinguish “calls transferred without location metadata”, “manual callback path active” and “normal live transfer restored”. A single green or red status hides the decisions people need to make.
Organisations that operate medical alert devices, private branch exchanges or large mobile fleets also need to understand how reconnection behaviour affects networks. The report does not assign them responsibility for Telstra's regulatory breaches, but it shows that device behaviour can become material system load. Coordinated testing can expose unexpected registration storms before they meet a safety-critical platform.
What the public record does not prove
The evidence does not prove that the medical alert devices were placing emergency calls during the registration surge. The final report says they were registering in preparation for possible calls.
It does not prove that all 494 affected calls were breaches. The finding comprised 127 transfer failures and 346 information-handoff failures. Twenty-one callers said assistance was not required, so the transfer duty did not apply to those calls in ACMA's conclusion.
It does not prove that every one of the 127 callers received no help. It establishes that their calls were not switched live to the emergency organisations and that Telstra used manual phone or email escalation.
It does not prove that the 346 transferred calls contained no caller information. The public number was available; the customer name and most precise available location were not provided.
It does not establish that the outage caused the death mentioned in Telstra's apology. The frozen public sources do not establish causation, and this article makes none.
It does not prove personal misconduct by a named worker. The useful accountability questions concern capacity, software recovery, contact ownership, fallback design, monitoring and organisational assurance.
It does not prove that announced remediation has closed every risk. Telstra described concrete actions, while independent operational evidence would be required to verify continuing effectiveness.
These boundaries do not weaken the article. They keep the conclusion tied to what the running service and public findings actually show.
The durable lesson
Emergency-call resilience cannot be reduced to whether a second database exists or whether staff have a list of phone numbers. Continuity is the successful journey experienced by the caller.
On 1 March 2024, the main information platform stopped responding. The separate contact list enabled many transfers, but eight destinations were wrong. The manual workaround passed information for callbacks, but it did not preserve 127 live calls. The 346 calls that were transferred did not carry the customer name and most precise location available to the platform.
Each stage recorded something useful. The primary system had caller data and destination records. The secondary system had backup contacts. Supervisors recorded details for escalation. None of those records alone was sovereign over the outcome. The reality layer was whether the person remained connected to the correct emergency organisation with the information that service already held.
That is a demanding standard, but it is understandable. A working fallback must preserve the call, the destination and the context. It must be tested under the conditions that disable the primary path. Its records must be reconciled with reachable endpoints. Its limits must be visible rather than hidden behind a reassuring label.
The goal is not to eliminate change or pretend that complex systems never fail. It is to ensure that failure does not force call-takers to discover the backup while a person waits for help.
Sources
- ACMA: Telstra pays $3 million penalty for Triple Zero outage
- ACMA final investigation report: Telstra ECP disruption, 1 March 2024
- Telstra: Our Triple Zero outage — the facts, the cause, and what's next
- Federal Register of Legislation: Emergency Call Service Determination version in force during the incident
Member Briefing
Deeper Profile Context
Sign in with the right membership level to unlock the full briefing and source notes.
Only for Strategic Circle
Strategic Circle
Open to all readers. Unlock profile briefings after joining and signing in.
Join Strategic CircleOnly for Leadership Alliance
Leadership Alliance
For qualified IP-asset owners and management; sign in to unlock alliance briefings.
Join Leadership Alliance
