On June 12, 2025, Cloudflare experienced a service outage lasting 2 hours and 28 minutes that affected a broad set of critical products. The failure originated in the storage infrastructure underlying Workers KV, which serves as a dependency for configuration, authentication, and asset delivery across many Cloudflare services. Part of that infrastructure is backed by a third-party cloud provider that suffered its own outage the same day, directly impacting KV availability. Cloudflare stated that the proximate cause was a vendor failure but that it bears ultimate responsibility for its dependencies and how it architects around them.

The incident was not caused by an attack or other security event, and no data was lost. Magic Transit and Magic WAN, DNS, Cache, proxy, WAF, and related services were not directly affected.

BLOG-2847 hero image

How Workers KV failed

Workers KV is designed as a "coreless" service: it runs independently in each Cloudflare location, so there should be no single point of failure. It nonetheless depends on a central data store as the source of truth. Failure of that store produced a complete outage for cold reads and writes to the KV namespaces used across Cloudflare's products.

Workers KV is being moved to more resilient infrastructure for its central store, including a migration to Cloudflare R2, intended to eliminate data consistency problems from the original syncing architecture and to better support data residency requirements. As part of that re-architecture, a storage provider was removed, leaving a coverage gap that this incident exposed.

Why the blast radius was wide

Cloudflare's stated principle is to build its services on its own platform building blocks, and Workers KV is a major one. Under normal conditions this gives internal and external products robust shared storage rather than each team building its own. When KV failed, the cascading effects significantly broadened the incident's impact.

Time

Event

2025-06-12 17:52

INCIDENT START
Cloudflare WARP team begins to see registrations of new devices fail and begin to investigate these failures and declares an incident.

2025-06-12 18:05

Cloudflare Access team received an alert due to a rapid increase in error rates.Service Level Objectives for multiple services drop below targets and trigger alerts across those teams.

2025-06-12 18:06

Multiple service-specific incidents are combined into a single incident as we identify a shared cause (Workers KV unavailability). Incident priority upgraded to P1.

2025-06-12 18:21

Incident priority upgraded to P0 from P1 as severity of impact becomes clear.

2025-06-12 18:43

Cloudflare Access begins exploring options to remove Workers KV dependency by migrating to a different backing datastore with the Workers KV engineering team. This was proactive in the event the storage infrastructure continued to be down.

2025-06-12 19:09

Zero Trust Gateway began working to remove dependencies on Workers KV by gracefully degrading rules that referenced Identity or Device Posture state.

2025-06-12 19:32

Access and Device Posture force drop identity and device posture requests to shed load on Workers KV until third-party service comes back online.

2025-06-12 19:45

Cloudflare teams continue to work on a path to deploying a Workers KV release against an alternative backing datastore and having critical services write configuration data to that store.

2025-06-12 20:23

Services begin to recover as storage infrastructure begins to recover. We continue to see a non-negligible error rate and infrastructure rate limits due to the influx of services repopulating caches.

2025-06-12 20:25

Access and Device Posture restore calling Workers KV as third-party service is restored.

2025-06-12 20:28

IMPACT END 
Service Level Objectives return to pre-incident level. Cloudflare teams continue to monitor systems to ensure services do not degrade as dependent systems recover.

INCIDENT END
Cloudflare team see all affected services return to normal function. Service level objective alerts are recovered.

BLOG-2847 2

Workers KV error rates to storage infrastructure. 91% of requests to KV failed during the incident window.

BLOG-2847 1

Cloudflare Access percentage of successful requests. Cloudflare Access relies directly on Workers KV and serves as a good proxy to measure Workers KV availability over time.

Workers KV, WARP, Access, Gateway, Images, Stream, Workers AI, Turnstile and Challenges, AutoRAG, Zaraz, and parts of the Cloudflare Dashboard were impacted, affecting Cloudflare customers globally. All timestamps in the timeline above are in Coordinated Universal Time (UTC).

Remediation

Cloudflare is accelerating planned resiliency work, organized into several active workstreams aimed at removing singular dependencies on storage it does not own and at improving recoverability for critical services including Access, Gateway, and WARP:

  • Improving redundancy inside Workers KV's storage infrastructure to remove dependency on any single provider. During the incident window, teams began cutting over and backfilling critical KV namespaces to Cloudflare's own infrastructure in case the outage continued.
  • Short-term blast radius reductions for individual affected products, so each becomes resilient to the loss of any single point of failure, including third-party dependencies.
  • Tooling for progressively re-enabling namespaces during storage infrastructure incidents, so key dependencies such as Access and WARP can come up without risking a denial-of-service against Cloudflare's own infrastructure while caches repopulate.

The list is not exhaustive; teams continue to revisit design decisions and assess near- and long-term infrastructure changes to mitigate similar incidents. Cloudflare said it is sorry for the impact on organizations of all sizes that depend on it.