Summary
- DigitalOcean marked a critical incident in its Kansas City
MKC1region resolved at 19:34 UTC on 19 August after restoring the regional control plane and broad CPU/platform services. - The same closing update said GPU Droplets remained impacted and all nodes were not yet restored, placing regional recovery and residual GPU recovery on separate public and customer-specific tracks.
A status page can close an incident before every workload has returned. DigitalOcean made that distinction unusually explicit on 19 August: it marked the MKC1 regional interruption resolved while saying in the same update that GPU Droplets continued to be impacted and that engineers were still restoring nodes.
The incident began at 11:10:06 UTC, when the company reported an issue across multiple racks and nodes in Kansas City. It warned that GPU workloads could be disrupted and DigitalOcean Kubernetes, or DOKS, worker nodes could enter the NotReady state. The public incident was classified critical.
At 12:22, DigitalOcean said it had identified the root cause and was applying remediation. It did not publish what the cause was. The affected surface named in that update had widened: GPU workloads and Serverless Inference could be disrupted, DOKS workers could remain NotReady, and customers with affected clusters might be unable to reach Kubernetes API endpoints or perform cluster-management operations. By 15:28 the provider said it was restoring connectivity and bringing affected nodes online, without quantifying the racks, nodes, clusters or customers involved.
The decisive split appeared at 18:59. DigitalOcean said the regional control plane was fully healthy. CPU Droplets, Managed Databases, Load Balancers, Block Storage and Spaces were operating normally; Droplet creation, resizing and other management actions were working; DOKS control planes were reachable. Those are meaningful recovery claims across shared regional and control services.
But the update immediately drew a second boundary. GPU Droplets in MKC1 remained offline or unreachable. DOKS GPU worker nodes could remain NotReady, and GPU-backed inference endpoints in the region could remain unavailable. DigitalOcean said it would monitor the regional control plane briefly, resolve the public incident and give GPU customers personalised information through Slack and email instead of the status page.
At 19:34:54, it did so. The final update again described the regional control plane, CPU Droplets and the named platform services as healthy or normal, yet also said GPU Droplets continued to be impacted and teams were working to restore all nodes. The 8-hour, 24-minute public incident window is therefore a status-management interval with a changing affected-service set. It is not evidence that every service was unavailable throughout, nor that every GPU workload recovered when the label changed to resolved.
The timing matters because MKC1 was new. DigitalOcean announced the Kansas City region on 4 August, 15 days before the incident. Its launch notice described a fully liquid-cooled facility running NVIDIA B300 GPUs and advertised a broad colocated product surface: Droplets, DOKS, databases, object and block storage, load balancing, networking, App Platform and Functions among other services. The incident does not establish that liquid cooling or B300 hardware caused the failure. It does make recovery evidence in a newly commissioned GPU region commercially relevant.
DOKS also illustrates why layer-specific language matters. DigitalOcean manages the Kubernetes control plane, while customer workloads run on worker nodes. A reachable control plane lets an operator inspect and manage a cluster; it does not create accelerator capacity on a GPU worker that is still NotReady. DigitalOcean documents an automatic worker-node remediation feature, but it is a public-preview, opt-in control with configured conditions, grace periods, actions and in-flight budgets. Nothing in the incident record says affected customers enabled it or that it repaired these nodes.
The public record stops short of a full post-incident account. DigitalOcean said a cause had been identified and later referred to working with the facility, but it disclosed no causal mechanics. It published no customer denominator, no exact time when all GPU nodes returned, no data-loss or workload-integrity finding, and no detailed prevention plan. Later green component metadata cannot supply those missing workload facts retrospectively.
For cloud operators, the practical lesson is not to distrust status pages; it is to read their scope. Regional control-plane health, CPU product availability, worker-node readiness and GPU endpoint reachability are separate measurements. DigitalOcean's own sequence preserves that distinction. The operational question after 19:34 was no longer whether the shared MKC1 region could accept normal control operations. It was which GPU resources remained impaired, what customer failover was available and when those workloads—not merely the incident record—were actually ready.
Sources
- https://status.digitalocean.com/incidents/qd8v4wddyqjr
- https://status.digitalocean.com/api/v2/incidents/qd8v4wddyqjr.json
- https://ideas.digitalocean.com/changelog/now-available-kansas-city-data-center
- https://docs.digitalocean.com/platform/regional-availability/
- https://docs.digitalocean.com/platform/
- https://docs.digitalocean.com/products/kubernetes/details/managed/
- https://docs.digitalocean.com/products/kubernetes/how-to/configure-automatic-node-remediation/
- https://docs.digitalocean.com/products/kubernetes/details/availability/
Member Briefing
Deeper Profile Context
Sign in with the right membership level to unlock the full briefing and source notes.
Only for Strategic Circle
Strategic Circle
Open to all readers. Unlock profile briefings after joining and signing in.
Join Strategic CircleOnly for Leadership Alliance
Leadership Alliance
For qualified IP-asset owners and management; sign in to unlock alliance briefings.
Join Leadership Alliance

