Anatomy of a Multi-Day Control Plane Outage
From November 2, 2023, at 11:43 UTC through November 4 at 04:25 UTC, Cloudflare's control plane and analytics services were down or degraded. The control plane—the customer-facing web interface and APIs for all Cloudflare products—failed. Analytics and raw log services were unavailable for most customers for the duration. The company restored the majority of control plane services at its European disaster recovery facility by November 2 at 17:57 UTC, but full restoration of all services, including those dependent on the failed data center, took until November 4.
Network and security services continued to pass traffic throughout the incident. Customers, however, experienced periods where they could not make configuration changes to those services.
The Architecture That Was Supposed to Prevent This
Cloudflare designed its control plane and analytics systems to run across three independent data centers near Hillsboro, Oregon. The facilities are spaced far enough apart to avoid common natural disasters but close enough to operate as continuously syncing, active-active redundant clusters. With this design, any single facility going offline should leave the remaining two able to continue running.
The high availability architecture was a four-year rollout. While most critical systems had been migrated, some services—especially newer products—had not yet been onboarded. Logging systems were deliberately excluded from this cluster. The reasoning was that logs could queue at the network edge until the logging facility returned, making delayed analytics an acceptable trade-off.
Power Failure at PDX-04
The largest of the three facilities, operated by Flexential and referred to as PDX-04, housed Cloudflare's primary analytics cluster, more than a third of the high availability cluster machines, and the default location for services not yet integrated into that cluster. Cloudflare consumes roughly 10 percent of the facility's total capacity.
At 08:50 UTC on November 2, Portland General Electric experienced an unplanned maintenance event on one power feed into PDX-04. Flexential compensated by starting generators while keeping the second utility feed active. Critically, Flexential did not inform Cloudflare of the generator failover—no observability tool could detect a change in power source. Had Cloudflare known, it would have stood up monitoring and proactively migrated control plane services out of the degraded facility.
The dual-source operation was itself unusual. Standard practice might have been to run exclusively on generators or solely on the remaining utility feed. Cloudflare has not received a clear explanation for why Flexential operated both concurrently. One hypothesis involves a Portland General Electric program (DSG) where the utility runs data center generators to support the grid in exchange for maintenance and fuel. Flexential has not confirmed whether DSG was active.
Catastrophic Failure Chain
At approximately 11:40 UTC, a ground fault occurred on a PGE transformer at PDX-04. The fault likely tripped protective measures that shut down not only the remaining utility feed but also all 10 of the facility's generators. The UPS battery bank, designed to provide roughly 10 minutes of bridging power, began failing after only 4 minutes, based on Cloudflare's own equipment monitoring. Generator restoration took longer than the batteries could bridge.
Flexential's internal recovery was hampered by significant operational issues. According to reports from Flexential employees, three factors delayed generator restart: the ground fault had tripped circuits requiring physical access and manual restart, Flexential's access control system was not on battery backup and was offline, and the overnight shift consisted of a security guard and a technician with only one week of experience—no experienced electrical or operations expert on site.
All customers in the facility lost power between 11:44 and 12:01 UTC. Flexential never proactively notified Cloudflare of the problem. Cloudflare detected the outage when two facility routers went offline at 11:44 UTC and dispatched its own team to the site. Flexential's first communication—a message saying engineers were working to restore power—came at 12:28 UTC.
Failures in the High Availability Design
Cloudflare had planned for a full data center loss, but testing had not covered the complete scenario. The high availability cluster had been tested by taking each of the other two data centers entirely offline, and by taking the high availability portion of PDX-04 offline. It had never been tested with full removal of the entire PDX-04.
That gap proved significant. A subset of services running in the high availability cluster had dependencies on services exclusively located in PDX-04. In particular, Kafka and ClickHouse—the log processing and analytics backends—were only in PDX-04, yet services that depended on them were running in the high availability cluster. These dependencies should have been looser, should have failed more gracefully, and should have been caught in testing.
Cloudflare also acknowledged being too lax about requiring products to integrate with the high availability cluster before being declared generally available. New products often run on bespoke backends that vary by team. Best-practice migration existed as a culture but not as a formal requirement. As a result, redundancy protections were inconsistent by product.
More broadly, too many services depended on the availability of the core facilities. The distributed network performed as expected throughout, but configuration and management of too many products required favorable conditions in the core. Cloudflare's own infrastructure did not fully leverage the distributed systems it sells to customers.
Recovery and Disaster Failover
Flexential restored generators at 12:48 UTC and power gradually returned to the facility. But when technicians attempted to power up Cloudflare's circuits, the circuit breakers were discovered faulty. Whether the breakers failed due to the ground fault, a surge, or prior undetected damage is unknown. Replacement parts needed to be sourced as more breakers had failed than Flexential had on hand.
With more services offline than expected and no restoration timeline from Flexential, Cloudflare made the call at 13:40 UTC to activate disaster recovery sites in Europe. Only a small percentage of the overall control plane needed to fail over, as most services ran on the two surviving active data centers. Services began turning up at the disaster recovery site by 13:43 UTC.
The surge of previously failing API calls created a thundering herd problem, which was addressed with rate limits. Customers of most products saw intermittent errors during this period. By 17:57 UTC, services on the disaster recovery site stabilized. Some systems still required manual configuration (such as Magic WAN), and other services—particularly log processing and bespoke APIs—remained unavailable until PDX-04 came back.
Products That Took Longer
Several newer products had no fully implemented or tested disaster recovery plan and did not stand up on the recovery site properly. These included the Stream service for uploading new videos, among others. Teams worked two tracks simultaneously: reimplementing these services on the disaster recovery site and migrating them to the high availability cluster.
Flexential replaced the circuit breakers and restored both utility feeds with confirmed clean power at 22:48 UTC on November 2. Cloudflare's team decided to rest for the night and begin the migration back in the morning—deliberately trading full recovery speed for reduced risk of compounding errors.
The return to PDX-04 began November 3 with physically booting network equipment and powering up thousands of servers. The process was a full bootstrap; the state of services was unknown following multiple likely power cycles. Configuration management servers took 3 hours to rebuild, after which the rest of the servers were rebuilt in parallel, each taking 10 minutes to 2 hours, with some services requiring sequenced starts due to dependencies.
Full restoration completed November 4 at 04:25 UTC. For most analytics datasets, logs continued to be replicated to European core data centers during the outage, so there should be no data loss for those products. Some datasets not replicated in Europe have persistent gaps. Customers using log push should assume logs were not processed for the majority of the event and will not be recovered.
Remediation Plans
Cloudflare announced a company-wide shift in engineering priority modeled loosely on Google's Code Yellow/Code Red crisis allocation. The internal version, Code Orange, redirects non-critical engineering resources toward control plane reliability. Specific commitments include:
- Removing dependencies on core data centers for control plane configuration, pushing functionality to the distributed network wherever possible
- Ensuring the control plane continues to function with all core data centers fully offline
- Requiring all generally available products to use the high availability cluster without software dependencies on specific facilities
- Requiring tested, reliable disaster recovery plans for all generally available products
- Testing failure blast radius and minimizing the number of services impacted by any single failure
- Implementing more rigorous chaos testing, including full removal of each core facility
- Auditing all core data centers with plans for reauditing against compliance standards
- Developing a logging and analytics disaster recovery plan that guarantees no log loss, even with total core facility failure
Cloudflare's post-mortem is candid about the operational miss: the architecture was theoretically sound enough to survive even a cascading provider failure, but nondeterministic testing and lax enforcement of its own integrations created the conditions for a prolonged outage.



