Summary

  • ThousandEyes places DynamoDB DNS restoration at 09:25 UTC on October 20, 2025, but reports new EC2 launch failures or connectivity problems until 20:50 UTC. Those are different recovery milestones, not a uniform outage lasting between them. ThousandEyes analysis
  • Dipak Kr das’s technical account describes accumulated lease-recovery work and a separate network-update backlog. Restoring access to the initiating dependency therefore did not finish the downstream recovery. Technical account on Medium

A system can become reachable again while the systems depending on it still have unfinished work. That is the important mechanism in the October 19–20, 2025 Northern Virginia incident involving Amazon Web Services. In Dipak Kr das’s account, DynamoDB’s return allowed EC2’s management machinery to resume work, but the volume of that work became another obstacle to recovery. The problem was no longer only whether a dependency could answer; it was whether downstream systems could make progress. Incident explanation

A DNS milestone, not an application all-clear

Dipak Kr das attributes the initiating DynamoDB endpoint-resolution failure to a latent race condition in automated DNS management. This is his explanation of the incident, not an independently established reconstruction presented here. The narrower chronology used below comes from ThousandEyes, which distinguishes restoration of DNS information from the subsequent return of successful connections. DNS explanation

ThousandEyes reports these milestones for October 20, 2025, all in UTC:

Time Reported state What it does not establish
09:25 DynamoDB DNS information restored Every client could immediately reconnect
09:25–09:40 Successful endpoint resolution and connections returned as cached records expired Downstream provisioning had completed its recovery
Until 20:50 New EC2 launches failed or experienced connectivity issues Every launch failed continuously until that time

The distinction is material: the final entry combines launch and connectivity impairment rather than identifying one uninterrupted failure affecting every request. ThousandEyes chronology

Subtracting 09:25 from 20:50 produces an interval of 11 hours 25 minutes. That measures the distance between the selected DNS-restoration milestone and the reported end of launch-or-connectivity impairment. It is not a customer outage duration, an application recovery measurement or an estimate of lost work. A single duration would conceal precisely the changing conditions that engineers needed to distinguish.

Recovery itself had work to complete

EC2’s Droplet Workflow Manager, or DWFM, is part of the management machinery for the physical servers hosting instances. ThousandEyes reports that DynamoDB unavailability prevented it from completing required state checks, disrupting lease management. Here, a lease means an internal, time-bounded management relationship—not the customer’s commercial rental agreement for a virtual machine. State-check dependency

Dipak Kr das’s account explains the next constraint. After DynamoDB returned, a surge of lease re-establishment work overwhelmed DWFM and prevented forward progress. Engineers responded by throttling incoming work and selectively restarting DWFM hosts. That describes a recovery problem caused by accumulated work, not merely a dependency that remained unreachable. Lease-recovery explanation

This distinction changes how to read the interventions. Limiting incoming work can help a congested system complete what is already pending. Selectively restarting an internal manager is also different from restarting customers’ virtual machines. The account concerns AWS’s internal management systems; it should not be turned into an instruction for customers to reproduce those actions on their own instances.

A second boundary remained after launch capability began to return. The same account describes a Network Manager backlog in propagating network state to newly launched instances. Some new instances consequently lacked connectivity or failed health checks. A launch could therefore advance further without yet delivering a machine ready to serve an application. This was reported as internal network-state propagation, not an Internet-wide routing failure. Network-update backlog

The sequence has several separate completion conditions: reach the dependency, restore management progress, finish the necessary network setup, and complete a useful application operation. Success at one condition is evidence about that condition. It is not proof that the others have been satisfied.

What surviving instances actually protected

Gremlin’s reliability account distinguishes instances started before the event, which it describes as staying healthy, from customers’ difficulties starting new instances after DynamoDB recovered. That is an important protection, but a limited one: continued instance execution does not establish end-to-end application availability. An application can still require a service or recovery operation outside the surviving machine. Gremlin’s reliability analysis

The evidence therefore supports separating existing execution from replacement or expansion capacity. It does not support saying that every workload failed, or that every surviving instance continued to deliver a healthy customer application. Nor do these accounts identify which particular customer could bypass a failed dependency through another zone or Region.

For a recovery design, the useful question is conditional: does keeping service running require creating something new while provisioning remains impaired? Existing capacity and capacity that must still be created are not interchangeable assurances. That does not make geographical redundancy useless; it makes the operations required to activate it part of the design that needs testing.

The limits of the reconstruction

This retrospective draws on technical analyses by ThousandEyes, Dipak Kr das and Gremlin, not an independent forensic audit. The detailed internal recovery mechanism remains an attributed explanation. These accounts do not establish a common recovery time for every AWS service or customer application, and they do not establish the effectiveness of today’s remediation.

The bounded lesson is nevertheless useful. Repairing the first failed dependency and restoring usable new capacity are different achievements. A recovery declaration should say which one has actually been demonstrated.