Summary

  • Cloudflare says its control plane and analytics services were disrupted beginning on November 2, 2023, while core edge traffic continued to flow [1].
  • The disclosed trigger was a power failure sequence at PDX-04, a data center in Oregon that hosted critical control-plane dependencies [1][3].
  • The incident exposed hidden dependencies in Kafka, ClickHouse, authentication, internal tooling, and other systems needed to recover or observe parts of the service [1].
  • A second power failure at the same facility in 2024 became a live test of Cloudflare's Code Orange preparation and reduced the control-plane impact [2].
  • A credible closeout must prove that configuration, analytics, identity, internal tooling, and customer-visible control functions can survive a full facility loss, not merely that the edge keeps serving cached or already-configured traffic.

What happened

Cloudflare's November 2023 postmortem says the outage began at 11:43 UTC on November 2. The company described the affected layer as the control plane and analytics services. That means customers and operators could face trouble changing settings, seeing analytics, or using administrative features even when the distributed edge continued to process traffic that did not need a new control-plane decision [1].

The physical starting point was PDX-04 in Oregon. Cloudflare says Portland General Electric had an unplanned maintenance event affecting one independent power feed into the building. The company then described a sequence involving utility feeds, batteries, generators, building access, replacement of circuit breakers, and staged restoration of servers [1]. Baxtel's public report also identifies the disruption as tied to a power failure at a Flexential data center in Hillsboro, Oregon, and quotes the same Cloudflare account about the utility-feed event [3].

The important reader-level point is simple: the Internet edge can be alive while the control layer is impaired. Existing routes, caches, and already-deployed configuration may continue to work. But a customer who needs to change policy, inspect analytics, adjust security controls, or confirm recovery depends on systems behind the edge. When those systems sit behind hidden dependencies in one facility, the service has a continuity gap.

The hidden dependency was not one machine

The postmortem does not describe a single server failure. It describes an ecosystem of dependencies. Cloudflare names control-plane and analytics components, internal operational tooling, Kafka, ClickHouse, identity and authorization dependencies, and the systems needed to restart or rebuild services [1]. Some dependencies were understood. Others were not fully visible until the facility failed.

This is why the incident belongs in network-infrastructure accountability rather than generic cloud outage coverage. A control plane is a network operator's authority surface. It is where a change becomes running configuration, where an alert becomes operational action, and where a customer sees whether a security or routing state is healthy. If that layer depends on one facility in ways that are not tested under full loss, the public risk is not just downtime. It is loss of operational visibility and change authority.

The postmortem also separates what kept working from what did not. Cloudflare says its network continued to serve most customer traffic. That should not be reduced to "the outage was minor." It shows a split-brain accountability problem: the data plane can look resilient while the control plane is brittle. A mature post-incident test has to inspect both layers separately.

Disaster recovery is a running-code test

Cloudflare says it had disaster recovery plans, but the event showed that some services were not yet ready for the total loss of the facility and that some procedures had not been tested under the right conditions [1]. A plan on paper can name a backup location. Running code decides whether configuration, credentials, data replication, queues, dashboards, and internal tools actually restart in time.

The recovery also had a demand problem. When services started returning, users and internal systems could create a thundering herd: many clients reconnecting, refreshing, or retrying at once. That is a control-plane capacity question. It is not enough for a standby system to start. It must absorb the backlog, preserve correctness, and avoid turning recovery into a second outage [1].

The 2024 follow-up is useful because it provides a public comparison. Cloudflare says the same facility had another major power failure four months later. This time, the company activated Code Orange, a prepared internal incident posture, and reported that prior changes reduced the customer impact [2]. That does not prove every dependency is solved. It does show the right evidence pattern: a similar failure occurred, the operator had made changes, and the second outcome was materially different.

Physical ownership still matters in a cloud service

A common mistake is to treat cloud control planes as abstract software. The PDX incident shows that physical ownership still matters. The public record includes utility feeds, generator behavior, battery depletion, building access, facility communication, circuit breakers, and server restart order [1][3]. Those details sit below the API but above the customer outcome.

Third-party facility dependency is not automatically a failure. Large networks use colocation and cloud-like operating models because they need scale and reach. The accountability question is what the operator can prove when a facility owner, utility provider, contractor, or physical access rule becomes part of the recovery path. The dependency should be explicit enough that the customer-facing service can survive the dependency failing or becoming slow.

Baxtel's report matters here because it places the incident in the data-center operator context and includes a public Flexential response about utility scenarios and grid-support arrangements [3]. The article should not overstate that response into a full root-cause adjudication. It does show that Cloudflare's recovery chain crossed organizational boundaries, which is exactly where accountability records often become thin.

The doctrine surface is operational continuity

The Heng.lu surface is operator continuity and hosting/network identity. A directory or public status page can identify Cloudflare as the operator and PDX-04 as a facility dependency. Those records do not prove that control-plane state, credentials, analytics data, and internal tooling can survive a full facility loss. The reality layer is running recovery evidence.

That distinction keeps the article away from blame theater. The question is not whether Cloudflare, Flexential, or the utility should be rhetorically blamed in isolation. The question is which dependency was allowed to sit on the critical path, whether the operator knew that before the outage, and what test now proves that the same dependency cannot disable control authority again.

In network terms, the exposed surface is configuration continuity. Customers buy a service partly because they can change policy during abnormal conditions. A security rule, DNS setting, analytics query, tunnel control, or access-policy change may be most important during an incident. If the control plane is impaired just when customers need it most, resilience claims have to be more specific than aggregate edge capacity.

What a credible recurrence test would contain

First, simulate the full loss of the facility that hosts control-plane components. Do not merely power down one service. Remove facility-level assumptions: local network, storage, queueing, identity, dashboards, and physical access. Record which systems continue, which fail over, which enter degraded mode, and which need manual work.

Second, test customer-visible control paths while the failure is active. Can a customer change DNS or security configuration, view analytics, modify access policy, rotate credentials, or confirm status? If some functions are intentionally unavailable, the status page and product documentation should say so plainly.

Third, measure recovery under backlog. A standby control plane that works for quiet traffic may fail when every user and internal job retries at once. The test should include queue depth, reconnect rate, authentication load, database replication lag, dashboard latency, and error budget during the restart.

Fourth, keep the facility dependency record current. Name the third-party operator, utility feed assumptions, remote-hands path, circuit-breaker or generator dependencies where publicly disclosable, and the person or team authorized to activate an alternate plan. Confidential details can be protected, but the control owner and test result should not be invisible.

Finally, compare the next incident to the last one. The 2024 Code Orange report is a useful model because it ties a similar physical event to a changed response and a reduced impact [2]. Every later recurrence should be judged the same way: which prior lesson was tested, which control held, which new dependency appeared, and what customers could still do.

What to watch next

Watch for Cloudflare to keep separating edge data-plane health from control-plane and analytics health in public incident reporting. A single green status for global traffic is not enough if customers cannot change or observe configuration. Watch for language about full facility loss, not just individual service redundancy.

Also watch for whether Code Orange becomes a repeatable public evidence pattern. The March 2024 event suggests the company improved preparation after November 2023 [2]. The stronger signal would be a steady record of exercises and incidents showing that configuration, analytics, internal tools, and customer-facing control functions fail over independently of PDX-04 or any single facility.

The lasting lesson is that a cloud network can keep carrying traffic while its authority layer is impaired. Accountability means proving that the operator can still see, change, restart, and explain the service when a critical building goes dark.

Sources

  1. https://blog.cloudflare.com/post-mortem-on-cloudflare-control-plane-and-analytics-outage/
  2. https://blog.cloudflare.com/major-data-center-power-failure-again-cloudflare-code-orange-tested/
  3. https://baxtel.com/news/cloudflare-blames-flexential-dc-outage-for-its-service-disruption