Summary

  • At 11:50pm on 5 July 2024, Telstra mistakenly migrated the live production connection for Australia's 106 text emergency service instead of a test service. Connectivity returned at 12:36pm the next day.
  • The Australian Communications and Media Authority, or ACMA, found that the live route had not been identified as a mission-critical emergency service, had no emergency-traffic alarm, and was migrated without a check that confirmed the intended target.
  • The record identified no genuine 106 call during the incident. That fortunate outcome limits any claim about harm, but it does not make the control failure hypothetical: an external service operator discovered that a live emergency route had disappeared.

What 106 is, and why this was not an ordinary business line

In Australia, most people know Triple Zero, or 000, as the telephone number for police, fire and ambulance assistance. The number 106 serves a more specific access need. It is the national text emergency relay number for people who have a hearing or speech impairment and use a teletypewriter or textphone, commonly shortened to TTY.

A TTY is a device that lets a person type a conversation over a telephone connection. For a 106 emergency call, the caller types to a relay operator. The operator communicates the caller's information to the relevant emergency-service organisation and types the response back. This creates a chain rather than a single direct conversation: the caller, the text device, the telephone network, the 106 emergency call service, the relay operator and the police, fire or ambulance service all have to remain connected.

At the time of the incident, Concentrix Services operated the emergency call person for 106 through the National Relay Service. An “emergency call person” is the organisation formally responsible for receiving and handling a class of emergency calls. Telstra supplied Concentrix with a Session Initiation Protocol trunk, or SIP trunk, that supported the service.

SIP is a set of rules that communication systems use to begin, manage and end calls. A trunk is not necessarily one physical wire. It is a logical route through which many calls can pass between systems. A simple way to picture it is as a controlled bridge between Telstra's network and the relay operator's call platform. If the bridge disappears, the two sides may each remain powered on while the call journey between them no longer works.

This distinction matters because a communications line can appear small in an inventory while performing a critical public function. The 106 route did not carry the largest call volume in Telstra's network. It did carry an emergency path designed for a group that may not be able to use an ordinary voice call. Criticality therefore came from the consequence and uniqueness of the service, not from the number of packets or calls seen on an average day.

A production service was treated as a test service

ACMA's public investigation report gives the central event unusual clarity. At 11:50pm on 5 July 2024, Telstra was performing migration work. The intended target was a test service. Telstra instead migrated the production SIP trunk for 106 from one application server to another.

“Production” means the live environment serving real users. A test environment is meant for rehearsals and validation without exposing the live service. The difference is fundamental, but it is not established by the label alone. A system is production because real traffic and real obligations depend on it. If an inventory calls a route “test” while live emergency traffic is attached, the running state is authoritative and the inventory is wrong.

Telstra told ACMA that the production trunk had been incorrectly identified as the test service. Because the team believed it was moving a test service, the due diligence normally applied to production services did not occur. The pre-migration analysis was carried out on the test service. Limited validation during the migration did not alert the team that it was working on the live service. No validation was performed to confirm that the service selected for migration was the correct target.

That sequence turns a seemingly simple labelling error into a control-system failure. A label influenced the risk classification. The classification determined the depth of review. The review determined which checks were performed. The missing checks allowed a live emergency connection to be changed as if it had no public consequence.

The issue was not that a human being once used the wrong word. Every large network contains stale labels, inherited names and awkward identifiers. The accountability question is whether any one label can authorise a consequential change without independent evidence from the running network.

A safe migration gate would reconcile at least three views. The service inventory says what the asset is. The configuration and traffic records say what it is connected to and what it carries. An end-to-end probe says whether the critical user journey works. If those views disagree, the change should stop until the difference is explained. In this incident, the production identity did not survive that reconciliation because the reconciliation was not performed.

The timeline shows how the external boundary became the detector

The migration began at 11:50pm. The public report does not describe an immediate Telstra alarm showing that 106 had become unavailable. The next published milestone came from outside Telstra's own operating boundary.

At 8:12am on 6 July, Concentrix contacted Telstra. It had observed that it was not receiving 106 calls or test calls. Shortly afterwards, Concentrix also advised that it could not make outbound 000 calls from its fixed lines. The relay operator was not merely a customer reporting a slow application. It was the next operator in a public-safety chain reporting that the handoff no longer functioned.

Concentrix's ability to make 000 calls was restored at 10:47am. Connectivity to the 106 number was restored at 12:36pm. From the start of the migration to restoration of 106, the service had been unavailable for 12 hours and 46 minutes.

Telstra notified ACMA on 6 July. ACMA began a formal investigation on 20 August 2024 and later considered information supplied by Telstra at several stages. The regulator's report does not publish every internal ticket, alarm or recovery action, so it would be wrong to invent a detailed minute-by-minute engineering story. The public sequence nevertheless establishes the important boundary: the external relay operator detected the absence of calls and raised it with the carrier.

That is a weak position for a critical provider even when the downstream partner responds well. The provider changing the route should be able to see whether the route remains healthy. A downstream operator should be a second line of observation, not the primary alarm for a change made upstream.

In ordinary language, imagine that a building manager moves a door-control system after reading a tag that says “training door”. Eight hours later, the organisation on the other side reports that the only accessible emergency entrance no longer opens. The other organisation's report is useful, but it should not be the first reliable proof that the door was live.

The call path was short enough to understand, but long enough to fail at a handoff

The 106 path can be explained without treating every reader as a telecom engineer.

First, a caller uses a TTY or textphone to dial 106. Second, the telephone network recognises that number as an emergency service and carries the call toward the designated 106 emergency call person. Third, Telstra's SIP trunk connects the call into the Concentrix-operated relay platform. Fourth, a relay operator handles the text conversation and contacts the relevant emergency-service organisation on the caller's behalf. Fifth, information and responses travel back along the chain.

The trunk at issue sat at an organisational and technical boundary. Telstra controlled the carrier connection. Concentrix operated the relay service. Emergency-service organisations received the relayed request. This makes correct handoff records especially important. Each party needs to know which route is live, how to test it, whom to notify before a change, which symptoms indicate failure and who has authority to restore or reroute it.

ACMA reported that Telstra did not contact Concentrix before the migration began. That fact does not prove that every routine carrier change requires the partner's permission. It shows that the downstream operator of this emergency path did not receive a pre-change contact for the activity that disabled its connection.

The distinction between permission and coordination matters. Operational accountability does not require turning every network change into a political approval ceremony. It requires an accurate transfer of state. A partner responsible for the next stage of an emergency service may need to know that a route is being changed, what test window applies, what signal confirms success and how to raise a failure immediately. That record supports continuity without confusing the partner's role with ownership of the carrier's infrastructure.

Why “no genuine calls” is important evidence, not a reason to dismiss the event

The investigation found no genuine call to 106 during the incident. Telstra and Concentrix made test calls. This is the most important limiting fact in the public record.

It means the available evidence does not support a claim that a real caller attempted to reach 106 and failed, that emergency assistance was delayed, or that anyone was harmed. A responsible article should say that plainly. The absence of an identified genuine call is not a footnote to be buried beneath alarming language.

It is also not proof that the service could safely be unavailable. Emergency systems are maintained for unpredictable moments. Their value is precisely that nobody can schedule when a caller will need them. The control failure existed because a live, regulated emergency route was disabled, whether or not demand happened to arrive during that window.

This is a useful distinction between consequence evidence and exposure evidence. Consequence evidence asks what happened to actual callers. Here, the public record identifies no genuine call. Exposure evidence asks whether the service was capable of fulfilling its duty during the window. Here, the production route was unavailable for 12 hours and 46 minutes.

Both should be reported. Exaggerating the consequence would be inaccurate. Ignoring the exposure would make continuity depend on luck. The regulator's approach reflects that balance: ACMA described the potential seriousness, recorded the absence of a genuine call and still found a breach of the rule requiring proper and effective functioning of networks and facilities used for emergency calls.

For operators, near misses are valuable only when their limits and mechanisms are preserved. “Nobody called” should close the question of observed caller harm. It should open the question of why a live critical path could disappear without the upstream operator recognising it promptly.

A missing mission-critical flag changed the treatment of a real service

Telstra advised ACMA that the SIP trunk had not been identified in its IT systems as a mission-critical emergency service. As a result, migration work did not receive the more rigorous change process associated with that classification.

An inventory is often described as a database of equipment and software. For continuity work, it is better understood as a ledger of operational responsibility. A useful record does not merely say that a trunk exists. It says what the trunk carries, which public obligation depends on it, which organisation receives the handoff, which changes can affect it, which tests prove it works and who owns its recovery.

When the mission-critical attribute is missing, downstream controls may all behave “correctly” according to bad input. The change system sees a test asset. The approval workflow asks fewer questions. The maintenance window does not trigger a partner notice. Monitoring applies ordinary thresholds. The recovery team lacks a clear service map. Each step follows the record, while the live network follows reality.

This is why record accuracy is an engineering control rather than clerical hygiene. A stale asset name can be annoying. A stale criticality classification can change which safeguards run. A missing owner can delay escalation. An incomplete dependency map can cause a team to restore the server while leaving the call path broken.

The answer is not to declare every service mission critical. If everything receives the highest classification, the classification stops helping. The answer is to link criticality to verifiable service functions. Emergency routes, including low-volume accessibility paths, need a specific identity and an auditable reason for their classification. Periodic review should compare those records with active configurations and traffic.

A useful inventory check asks a plain question: if this object were disconnected now, which real user journey would stop? If the answer is an emergency service, the operating controls should not depend on the object having a familiar product name.

Alarms should follow emergency traffic, not the name on the object

ACMA also reported that Telstra had not configured alarms for SIP trunks carrying emergency traffic to improve visibility of adverse effects.

An alarm is a rule that turns a technical observation into an operator action. It may watch connection state, call attempts, failed responses, traffic absence, signalling errors or an end-to-end test. A good alarm is not simply loud. It detects the service condition that matters and reaches a person who can act.

For a low-volume service, monitoring requires care. An absence of real calls may be normal at many hours of the day. A dashboard cannot assume that zero genuine calls means failure. This makes controlled test traffic and state monitoring important. A synthetic probe can follow the route without pretending to be a real emergency. Connection-state alarms can detect a trunk that has disappeared. Reconciliation can compare expected and observed number translations. A partner acknowledgement can confirm the handoff.

The public report does not specify which exact alarm Telstra should have used or how often a test should run. It establishes the narrower fact that emergency-traffic alarms were not configured on the relevant SIP trunks. Any proposed control should therefore be framed as an operational lesson, not as a claim about an unpublished internal design.

A layered approach would reduce dependence on a single signal. Before a change, identify the route from active configuration and recent traffic. During the change, monitor trunk state and signalling. Immediately after the change, perform a controlled end-to-end test with the relay operator. For a defined period, watch for deviations and keep rollback authority available. If any layer disagrees, treat the production state as unresolved.

Monitoring also needs ownership. An alarm that appears on a generic dashboard but has no acknowledged operator is only a record of failure. The service ledger should identify the team that receives the alert, the relay partner contact, the incident severity and the maximum time before escalation.

Recovery was complicated by messages and number-range translation

ACMA said the production migration left Concentrix unable to receive 106 calls or make outbound calls through the SIP trunk. It also recorded associated technical issues involving a misalignment of messaging and number-range translations between Telstra and Concentrix. Those issues further complicated the incident and affected attempts to fix it.

“Number-range translation” sounds abstract, but the basic idea is familiar. Communication systems often rewrite or map telephone numbers so that each network recognises the destination in the format it expects. If two connected systems apply different mappings, a call can be directed incorrectly or rejected even when the underlying connection is present.

Messaging in this context refers to the signalling exchanged by the connected systems to establish and manage calls. Restoring a server or trunk does not guarantee that both sides interpret those messages and numbers the same way. A service can therefore move from “disconnected” to “connected but still unable to complete the journey”.

This is why rollback must be verified at the user-journey level. Reversing a change returns configuration toward a previous state. It does not automatically prove that partner state, translations, routing and call handling are aligned. The recovery record should show which end-to-end tests passed, which number formats were checked and which party confirmed service restoration.

The incident illustrates a broader infrastructure truth: boundaries accumulate state. Telstra had a view of the trunk. Concentrix had a view of incoming and outgoing calls. Each side had messaging and number rules. Safe migration required those views to remain compatible. When they diverged, the repair was no longer a single-server task.

This is not an argument against modernisation. Old application servers and inherited trunks must be migrated. The accountability standard is that the operator knows the live identity, coordinates the boundary and proves the complete service after the move.

What ACMA found, and what the finding does not say

At the time of the incident, subsection 11(1) of the Telecommunications (Emergency Call Service) Determination 2019 required carriers and carriage service providers, as far as practicable, to maintain the proper and effective functioning of controlled networks and facilities used to carry emergency calls.

ACMA found that Telstra failed to take practicable steps to ensure the 106 SIP trunk was maintained with an appropriate change process, operational documentation and visibility of adverse impacts. The regulator found one contravention of subsection 11(1). It also found related contraventions of the statutory obligation to comply with the determination and of the carrier-licence condition linked to compliance with telecommunications law.

ACMA announced that Telstra paid AUD 18,780, which the regulator described as the maximum penalty it could impose in the circumstances. Telstra also gave a court-enforceable undertaking. According to ACMA's public summary, that undertaking included improving relevant change-management processes, engaging an independent reviewer to examine operational arrangements supporting reliable delivery of 106, implementing reasonable recommendations, developing staff training and reporting progress to ACMA.

An infringement notice is not the same as a judicial trial, and a court-enforceable undertaking is not proof that every promised improvement is already finished. The public documents establish the regulator's finding, the penalty and the commitments. They do not support speculation about personal fault, hidden intent or undisclosed harm.

The size of the penalty should also be reported in context. ACMA said it was the maximum available in the circumstances. Comparing the amount with revenue and concluding that the service was considered unimportant would go beyond the evidence. The legal framework set the available response; the operational significance comes from the service and the control findings.

The most useful value of the report is its specificity. It names the live-versus-test identification error, the missing target validation, the absent mission-critical classification, the missing emergency-traffic alarms, the lack of pre-migration contact and the recovery complication. Those are concrete control surfaces that can be examined without turning the article into either corporate advocacy or outrage copy.

Running service is the reality layer

The deepest lesson is simple: the service that is carrying real obligations is production even when a database says otherwise.

Infrastructure organisations need inventories, registries and change records. Those records create shared memory. They make ownership visible and allow automation to apply the right controls. But a registry is a ledger, not a sovereign that can change reality by naming it. It records the network; it does not make an active emergency path become a harmless test route.

The running network is the reality layer. Active configuration, observed traffic and completed user journeys show what the system is doing. Documentation and labels must be reconciled to that layer. When they conflict, the safe response is not to trust the most convenient field. It is to stop, investigate and correct the record before continuing.

This principle is especially important for number resources and call paths. Telephone numbers must remain unique, accurate and transferable. Routing metadata must connect a number to the correct service. Security and operational context must show who can change the route. Continuity requires the next operator to receive the call in the format and state it expects.

The 106 incident joined all of those concerns. A unique emergency number pointed to a live relay service. A SIP trunk carried the boundary. An inventory described the wrong environment. The change process followed the inventory. Monitoring did not independently reveal the loss. The downstream operator observed the failure. Recovery then had to realign signalling and number translations.

That is a more useful account than saying “a server migration went wrong”. The server move was the action. The accountability failure was the inability of the control system to preserve the live identity and continuity of the service through that action.

Accountability should attach to roles and controls, not an unnamed worker

The public report does not name the individual who selected the service, approved the change, performed the migration or responded to the incident. There is no basis for assigning personal blame.

It is still possible to describe organisational accountability precisely. The service owner was responsible for an accurate critical-service record. Change management was responsible for target verification and appropriate review. Network operations was responsible for monitoring the live path. The partner-management function was responsible for an effective contact and test plan with Concentrix. Incident command was responsible for coordinating restoration and preserving the timeline.

These may not be the exact internal team names used by Telstra. They are functional roles that a resilient operator must assign. Naming the functions avoids two errors. One is scapegoating a technician for a system that supplied bad identity and weak checks. The other is using the word “system” so broadly that nobody owns the decision.

A durable remediation should make the correct action easier and the unsafe action harder. Training can remind staff why 106 matters. Inventory repair can attach the emergency-service classification. Change tooling can require independent target evidence. Monitoring can alert on loss. A pre-change workflow can obtain partner readiness. A post-change gate can require an end-to-end test before the window closes.

The public undertaking points in this direction by combining process improvement, independent review, training and reporting. The article cannot verify the unpublished implementation. It can identify the evidence that would demonstrate improvement: current service records, control samples, successful failover or migration drills, alarm tests, partner acknowledgements and tracked closure of review recommendations.

A practical migration standard for low-volume critical services

The controls suggested by this incident are understandable without specialist jargon.

First, give every critical service a unique identity that survives server names and vendor changes. The record should name the user journey, not only the technical component. “106 text emergency relay production path” is more useful than an inherited trunk code whose meaning lives in someone's memory.

Second, prove the target from live evidence. Before migration, compare the planned object with active configuration, recent traffic, destination numbers and the downstream partner's record. If the change is intended for test, demonstrate that no production obligation depends on it.

Third, require two-person confirmation for destructive or disconnecting actions on emergency routes. The second person should review independent evidence, not merely repeat the same label.

Fourth, make criticality machine-readable. When a route carries emergency traffic, the change platform should automatically apply the higher control tier, require a rollback plan and open the partner-notification and validation tasks.

Fifth, monitor state and journey. Connection alarms should detect a missing trunk. Controlled test calls should prove that the route reaches the relay service. Partner acknowledgement should confirm that incoming and outgoing behaviour is normal.

Sixth, coordinate the boundary. The upstream carrier and downstream relay operator should share the maintenance window, test procedure, escalation contact and restoration criteria. Coordination records a handoff; it need not give either party vague authority over the other's entire network.

Seventh, verify number and message translation after the move. A connected trunk is not enough if each side interprets signalling or number formats differently.

Eighth, keep the rollback available until the complete journey is proven. Do not close the change because a server process is green.

Ninth, preserve timestamps automatically. The record should show when the route changed, when monitoring detected a deviation, when the partner reported failure, when each recovery stage completed and who accepted the service back into operation.

Tenth, review quiet critical services deliberately. Low demand can hide failure. A route that receives few calls needs stronger synthetic testing, not weaker attention.

These controls do not guarantee that no outage will ever occur. They make the live state harder to mistake, the failure faster to detect and the recovery easier to prove.

What enterprise and public-sector buyers can ask

The incident has lessons for organisations that depend on carriers, hosted call centres, accessibility platforms or emergency communications.

A buyer cannot inspect every internal configuration. It can ask how the provider identifies critical routes across inventory, change and monitoring systems. It can ask whether low-volume services receive end-to-end probes. It can request evidence from recent resilience tests, not a generic statement that the platform is redundant.

Contracts and operating agreements should identify the service boundary. Who owns the carrier path? Who operates the relay platform? Who notifies whom before planned work? Which party declares restoration? How are failed tests escalated outside normal business hours?

Accessibility services should not be treated as optional add-ons whose continuity follows only after mainstream voice service. A service may support fewer users while being the only workable channel for those users. Risk classification should reflect substitutability and consequence, not only volume.

Public-sector resilience teams can also use the event as a tabletop exercise. Give participants an inventory that labels a live route as test. Ask which independent signal catches the conflict. Remove ordinary traffic so the absence of calls looks normal. Then test whether the service owner, change team, monitoring team and downstream partner can reconstruct the path and stop the change.

The important output is not a perfect slide deck. It is a corrected service record, a tested alarm, a named owner and a repeatable handoff.

What the public record does not prove

The evidence has clear boundaries.

It does not prove that a genuine user called 106 and failed. ACMA said no genuine call was identified during the incident.

It does not prove that nobody in Australia needed emergency assistance during the window. It establishes the call record available to the investigation, not every person's circumstances.

It does not prove intent. Telstra described an incorrect identification of production as test. The record supports mistake and insufficient control, not deliberate disruption.

It does not disclose every internal system, employee, contractor, ticket or alarm. Detailed claims about a particular individual or software defect would be speculative.

It does not establish that all emergency services were unavailable. The incident concerned the 106 relay path and Concentrix's related outbound calling through the trunk. It should not be merged with a national loss of 000.

It is not the same event as Telstra's 1 March 2024 Triple Zero call-centre disruption, which involved backup transfer numbers, or the July 2026 network timing outage. Keeping dates and mechanisms separate is essential to fair reporting.

It does not prove that the undertaking's work is complete. Evidence of implementation would require later reporting or verified control results.

These limits strengthen the conclusion. The documented facts alone show that a live emergency route was misidentified, changed without production safeguards, left without relevant classification and alarms, detected by the downstream operator and restored after a prolonged window.

Why this small route matters to large networks

Modern communications networks contain thousands of services that do not resemble a consumer product. They include interconnects, signalling links, number translations, accessibility routes, monitoring feeds and partner gateways. Many are old. Some are quiet. All are vulnerable to becoming invisible when inventories, teams and platforms change.

The danger is greatest during migration because the organisation is deliberately altering identity and location. A service moves from one server, platform or provider to another. Names change. Owners change. Test and production objects may look similar. The work is often scheduled when traffic is low, which also reduces the natural signals of failure.

That combination makes running-code evidence essential. Operators need to know which configuration is active, which traffic is real and which user journey depends on it. A registry entry can guide the work, but it must be continuously corrected by observed reality.

The 106 route also shows why resilience is relational. Telstra could not define success only inside its network. Concentrix needed to receive calls and place outbound calls. Emergency organisations needed to receive the relayed request. Number and message translations had to agree across the boundary. The service existed in the handoff.

For non-specialists, the accountability test can be reduced to five questions. What live service was being changed? How did the operator prove it was the intended target? Which alarm showed that the user journey still worked? Which partner confirmed the handoff? Which record proved restoration?

ACMA's findings show gaps across those questions. The fortunate absence of a genuine call prevented a documented caller consequence. It did not repair the records or make the route less critical.

The durable lesson is therefore neither “never migrate emergency systems” nor “a penalty will prevent mistakes”. Networks must change, and human error will remain possible. The better standard is to ensure that one mistaken label cannot carry enough authority to bypass the reality of a live service.

When an emergency line is running, its identity should be visible in configuration, inventory, monitoring, ownership and partner records. When it moves, each layer should independently confirm the same target. When it returns, the complete caller journey should be proven. That is how a critical service remains more than a name in a database.

Sources