Outage Post-Mortem: What Took Down Cloudflare’s Dashboard and API

On April 15, 2020, Cloudflare’s Dashboard and API were offline for over four hours. The incident began at 1531 UTC and was not resolved until 1952 UTC. The root cause was not a DDoS attack, a software bug, or a hardware failure—it was a series of disconnected cables during routine maintenance.

The Incident

Technicians were performing planned maintenance at one of Cloudflare’s two core data centers. The task involved decommissioning a cabinet filled with old, inactive equipment. That cabinet, however, also housed a patch panel that carried all external connectivity to other Cloudflare data centers. Over a three-minute window, the technician removed the unused hardware along with the cables in that patch panel.

This particular facility runs Cloudflare’s main control plane and database. When connectivity was severed, the Dashboard and API went down immediately. The broader network continued to operate normally—customer websites and applications stayed online, as did Magic Transit, Cloudflare Access, and Spectrum. Security services like the Web Application Firewall were unaffected.

The outage made several operations impossible:

  • Logging into the Dashboard
  • Using the API
  • Making configuration changes, including DNS records
  • Purging cache
  • Running automated Load Balancing health checks
  • Creating or maintaining Argo Tunnel connections
  • Updating Cloudflare Workers
  • Transferring domains to Cloudflare Registrar
  • Accessing Cloudflare Logs and Analytics
  • Encoding videos on Cloudflare Stream
  • Logging edge service information, causing a gap in customer log data

No configuration data was lost. Customer data remained intact on-site, and although backups and off-site replicas existed, neither was needed.

Response and Recovery

Cloudflare’s response involved two simultaneous tracks: restoring physical connectivity and preparing a failover to the disaster recovery core. With most of the team working remotely due to the COVID-19 emergency, engineers coordinated from two virtual war rooms, one focused on each objective.

Internal monitoring systems were quickly failed over so the SRE team could maintain visibility across the entire network—spanning more than 200 cities—and keep the edge service running smoothly.

As the incident unfolded, engineers reassessed every 20 minutes whether to fail over the Dashboard and API to disaster recovery or continue pursuing connectivity restoration. The failover itself was feasible, but the team knew from prior testing that failing back would be complex. The risk of physical damage (such as from a natural disaster) would have made the decision straightforward, but with no such damage, the team weighed the trade-offs in real time.

Connectivity was restored gradually:

  • At 1944 UTC, a backup 10Gbps link came online
  • At 1951 UTC, the first of four large Internet links was restored
  • At 1952 UTC, the Dashboard and API became available
  • At 2016 UTC, the second of four links was back
  • At 2019 UTC, the third link was restored
  • At 2031 UTC, fully redundant connectivity was achieved

Preventive Measures

Cloudflare identified several corrective actions to reduce the risk of recurrence. The primary issue was architectural: external connectivity used diverse providers and led to diverse data centers, but all connections terminated at a single patch panel, creating one physical point of failure. The company plans to distribute that connectivity across multiple parts of the facility.

Documentation was another gap. After the cables were unplugged, technicians lost valuable time identifying which connections were critical. Clear labeling of cables and panels would allow anyone responding to an incident to quickly find what needs to be restored.

Finally, the process itself needs refinement. When issuing instructions to retire hardware, the communication should explicitly call out cabling that must not be touched.

A full internal post-mortem is underway to ensure the root causes are fully addressed.