Summary
- Twilio opened a major incident at 16:47:26 UTC on 3 August for ConversationRelay customers using Deepgram speech-to-text and described the affected users as hard down.
- At 17:39:18 UTC, Twilio narrowed the affected configuration to Deepgram nova-2 and nova-3; it said Deepgram flux and Google speech models were not impacted.
- The first notice said customers could consider failing over to Google only if they had validated that Google worked for their use cases.
- Twilio said its engineers worked closely with Deepgram, observed recovery at 18:45:29 UTC and continued monitoring.
- The incident was marked resolved at 19:30:59 UTC, making the public incident lifecycle 2 hours, 43 minutes and 33 seconds.
- Twilio disclosed no root cause, affected-customer denominator, geography, remediation steps, automatic failover mechanism, data loss or security compromise.
The failure sat inside a specific runtime choice
ConversationRelay connects a live conversation to speech recognition and the surrounding application logic. In this incident, Twilio did not report that every ConversationRelay configuration failed, nor that every Deepgram model was unavailable. Its later update named nova-2 and nova-3 and explicitly excluded Deepgram flux and Google speech models from the affected set.
That precision matters because “the voice platform was down” would erase the configuration boundary. An application using one of the named models could have lost the speech-to-text step on which the rest of its dialogue depended. Another application using an unaffected path could continue. The status record supplies no customer count, call count, minute volume or geographic distribution from which to estimate the relative size of those groups.
The phrase “hard down” is still material. It indicates that, for the affected configuration, Twilio was not merely describing elevated latency or a small accuracy change. Yet it remains the provider’s qualitative wording, not proof that all Twilio voice services, all Deepgram traffic or every customer failed. A useful incident account preserves both halves: narrow technical scope and serious effect within that scope.
An available model is not automatically an available substitute
Twilio’s first update offered a conditional route: customers who had validated Google for their use cases might consider failing over. The condition carries more operational information than the vendor name. It acknowledges that changing a speech model is not necessarily equivalent to moving identical packets between interchangeable servers.
A conversational application may depend on language support, endpointing, punctuation, interim results, timing, confidence fields, domain vocabulary or how partial transcripts drive downstream actions. Even if two providers expose similar interfaces, their output and failure behaviour can differ. Twilio did not compare those characteristics here, and the incident provides no basis for ranking Google, flux, nova-2 or nova-3 on accuracy, speed, price or capacity.
Validation is therefore an application property. A team needs to know that the alternate configuration accepts the required audio, returns fields the orchestration layer can process, behaves acceptably under live timing and does not break compliance or customer-experience assumptions. Without that work, an “alternative” is an architectural possibility, not a production recovery control.
The timeline separates detection, diagnosis, recovery and closure
The incident began at 16:47:26 UTC. Twilio initially said engineers were investigating and suggested the conditional Google option. At 17:39:18, it said engineering had identified the issue, named nova-2 and nova-3 as hard down and said teams were working closely with Deepgram. It did not disclose what had been identified.
At 18:45:29, Twilio reported observed recovery and moved to monitoring. Resolution followed at 19:30:59. From the initial notice to final closure, 9,813.164 seconds elapsed. The monitoring transition came roughly one hour and 58 minutes after opening, with another 45 minutes used to observe whether recovery held.
Those states answer different questions. “Identified” means the operator believed it had isolated the issue, not that the public received a root cause. “Monitoring” establishes observed recovery, not permanent prevention. “Resolved” records the provider’s conclusion that the named service was operating normally. Collapsing all four into one outage timestamp would hide both the recovery test and the missing explanation.
Provider diversity has to reach the application layer
An architecture diagram can show two model providers and still have no usable failover. Credentials may be missing, quotas may be too low, data-processing terms may differ, a region may not be enabled, or the application may parse only one provider’s response. A manual switch might also take longer than the business process can tolerate.
The incident exposes the control surface: model selection, routing logic, configuration, contracts and test evidence. Resilience requires that these elements move together. A secondary provider whose path has never been exercised under representative audio is inventory, not continuity. Conversely, a tested alternate path can reduce dependence even when the primary provider’s internal root cause remains unknown.
This does not imply that every application should switch automatically. Automatic failover can amplify a fault, change transcription behaviour without human review or consume scarce quota. The correct design depends on the consequence of silence, delay, degraded recognition and inconsistent output. What the source supports is narrower: Twilio itself tied its workaround advice to prior validation.
Recovery does not answer what happened inside interrupted sessions
The final update says the affected models were operating normally. It does not describe whether calls were dropped, audio was buffered, transcripts were incomplete, requests were retried or sessions recovered in place. No data loss was reported, but absence of such a report is not proof that every in-flight interaction completed as intended.
Applications should reconcile at their own business boundary. That might mean checking for conversations with missing transcript segments, actions that depended on an absent utterance, unusual escalation rates or sessions that ended during the incident. The relevant record is not merely whether the API now returns success; it is whether each customer process reached an acceptable end state.
The same boundary prevents overstatement. The public record does not show customer-data exposure, interception or malicious activity. It also does not disclose transcript retention or recovery behaviour. Post-incident reconciliation is prudent because the application owns the outcome, not because the status page proves a security or data-integrity event.
Root-cause ownership remains deliberately unresolved
Twilio said its engineers worked closely with Deepgram to investigate and remediate. That sentence establishes collaboration across a dependency boundary. It does not assign fault, identify the failing component or describe what changed. The cause could sit in a provider service, an integration layer, routing, configuration or another dependency; those are possible categories, not findings.
A substantive postmortem would explain the initiating condition, detection, affected request path, why flux and Google remained available, what restored nova-2 and nova-3, and what recurrence controls were added. It would also quantify customers or traffic without exposing sensitive data. None of that appears in the incident history.
Until such evidence exists, customers can improve only the controls they own: configuration inventory, health probes, model-specific telemetry, explicit switch authority, tested rollback, quota readiness and session reconciliation. They cannot tune around a public root cause that Twilio has not supplied.
The commercial question is the cost of a silent dependency
Speech recognition can look like a replaceable API line item, but in a live conversational product it becomes part of the transaction path. If transcription stops, intent detection, routing, authentication or agent assistance may also stop. The cost is shaped by the application’s tolerance for missing or delayed speech, not simply by the incident’s provider label.
This makes continuity a purchasing and product-design decision. Teams need to price the engineering required to qualify an alternate model, maintain compatible prompts and schemas, monitor divergence and rehearse a switch. They also need to decide when consistency is more valuable than immediate failover. A backup that changes customer outcomes unpredictably can be worse than a controlled degradation.
The incident does not settle that trade-off. It provides a concrete test: customers with a validated Google path had an option Twilio could name; customers without one did not receive an equivalent universal workaround. The difference was preparation carried in the application, not a capability invented after the outage began.
Sources
Member Briefing
Deeper Profile Context
Sign in with the right membership level to unlock the full briefing and source notes.
Only for Strategic Circle
Strategic Circle
Open to all readers. Unlock profile briefings after joining and signing in.
Join Strategic CircleOnly for Leadership Alliance
Leadership Alliance
For qualified IP-asset owners and management; sign in to unlock alliance briefings.
Join Leadership Alliance

