Producers and consumers in a traditional RPC architecture have to agree on both scale and timing. When producers outpace consumers, or a consumer or downstream service goes offline, events are lost — and the problem grows when several consumers need to process the same data independently. An ecommerce backend, for instance, may emit transaction events that both an analytics system and a fraud detection service must read.

Decoupling the two sides solves this: a service sits in the middle, absorbing writes while independent readers consume at their own pace. Cloudflare K2, now in public beta, is that service — a durable event streaming primitive on the Developer Platform. Events sent to a K2 stream are stored as an ordered log, and consumers can read them either by dividing reads across a set of consumers or by delivering every message to every consumer. It is fully serverless, scales to large data volumes, and retains data long-term, so extended consumer downtime does not cause loss.

Under the hood, K2 is a partitioned, durable log built on R2 object storage, which is what lets it scale to huge storage volumes. A getting-started guide is available for creating a first stream.

Why K2 exists, and why it runs on R2

K2 was originally built to fill a need of Cloudflare's own: a durable buffer at the edge, first intended as the ingestion layer for Basin Pipelines. Pipelines runs on a pull-based stream processing engine, which means another system must hold events before they are read, transformed, and written to R2. Since Cloudflare commits to never dropping events once accepted into the Pipelines Stream, that holding area must remain durable for potentially long periods.

The conventional answer would be Apache Kafka. But Pipelines runs on Cloudflare's edge, which spans a large number of servers across more than 335 cities, and that architecture often rules out running traditional distributed systems software such as Kafka. Stateful services in particular face constraints: relatively small machine slices, relatively ephemeral machines, and networking that frequently traverses the public Internet. The same infrastructure offers advantages, though — proximity to users worldwide and substantial capacity for horizontal scale.

For the durable buffering system that became K2, the team chose to build on an existing state primitive: R2. Object storage of that kind pairs very durable storage — 11 9s — with strongly consistent APIs. Pushing replication and consensus down to the storage layer makes the application layer radically simpler, cheaper, and faster. It also separates compute from storage so each scales on its own, which keeps large volumes of historical data inexpensive to hold.

Building a log without appends

R2, like other object stores, has no append operation, so a log cannot be written incrementally. K2 instead writes complete segment files large enough to justify the cost of each read and write. Writes accumulate in memory on an edge service, and after a short wait for data to arrive, all events are written as a single segment file. R2's atomic operations provide ordering and strictly incrementing offsets without a separate coordination service.

The trade-off is produce latency. Writing to object storage is slower than writing to local disk, and the local batch must fill before the write begins. In the initial release, that comes to roughly 1 second of produce latency at the 99th percentile of response times. A deeper look at K2's design is planned for a future technical deep dive.

Choosing between K2, Queues, and Pipelines

K2 is not the only asynchronous delivery primitive on Cloudflare. Queues is built around tracking individual units of expensive or slow work that must complete asynchronously — an image processing application enqueueing a user request, for example — and it supports item-level logic such as retries, delays, and dead-letter queues for failed attempts. It shares surface similarities with K2 Streams: both accept events, store them durably, and deliver them to consumers.

The difference is the design target. K2 aims at high-scale data movement, long-term retention, and fan-out consumption. Messages are produced and consumed in batches, which enables efficient processing but gives up message-level retries, and the batching is also why producer latency is higher than for Queues.

Basin Pipelines is a serverless ingestion service: JSON events sent to a Pipeline can be transformed and written to R2 or a Basin Catalog. Pipelines is the recommendation when events ultimately land in object storage or Iceberg tables; K2 is the recommendation for custom processing or delivery to other destinations.

Producing, subscribing, and consuming

Usage begins with a stream. An account can hold many streams, one per use case or event type, created through cf, Wrangler, the dashboard, or the API. For a product analytics workload, the first step is creating a stream with cf:

$ cf k2 streams create --name app_events --http-enabled

{
  "id": "d78b09ee1f50430e9ec92a8af92b0231",
  "name": "app_events",
  "retention_seconds": 604800,
  "endpoint": "https://d78b09ee1f50430e9ec92a8af92b0231.k2.cloudflarestorage.com",
  "http": {
    "enabled": true,
    "authentication": false
  },
  "worker_binding": {
    "enabled": true
  },
  "created_at": "2026-09-28T15:14:39.053Z",
  "modified_at": "2026-09-28T15:14:39.053Z"
}

With the stream in place, events can be produced through an HTTP API or a Worker binding — for example, from a Worker:

const result = await env.EVENTS.send([
  {
    content: new TextEncoder().encode(
      JSON.stringify({
        event: "page_view",
        path: new URL(request.url).pathname,
        timestamp: Date.now(),
      }),
    ),
    headers: { "content-type": "application/json" },
  },
]);

if (!result.success) {
  console.error(`Produce failed: ${result.error.message}`);
  return new Response("Failed to record event", {
    status: result.error.retryable ? 503 : 500,
  });
}

K2 treats data as bytes, so applications are free to choose whatever format or encoding suits them.

Reading requires a subscription. Subscriptions divide work among consumers to provide read parallelism, scaling to more readers than a single server could support. One can be created through the HTTP API:

$ curl -X POST "https://d78b09ee1f50430e9ec92a8af92b0231.k2.cloudflarestorage.com/subscriptions" \
    -H "Authorization: Bearer ${CLOUDFLARE_API_TOKEN}" \
    -H "Content-Type: application/json" \
    --data '{
      "name": "analytics_processor",
      "start_at": { "type": "earliest" }
    }'
{
  "result": { "id": "ee13f761783d3823a447a47b572ebf76" },
  "success": true,
  "errors": [],
  "messages": []
}

Each consumer then polls the subscription:

$ curl -X POST \"https://4d8f5394e3e733debdeca9c65c5b7439.k2.cloudflarestorage.com/subscriptions/ee13f761783d3823a447a47b572ebf76/consume" \
    -H "Authorization: Bearer ${CLOUDFLARE_API_TOKEN}" \
    -H "Content-Type: application/json" \
    --data '{
      "worker_id": "analytics-1",
      "max_records": 100
    }' 
{
  "result": {
    "batch_id": "b7e4c9210a3f468d95c2e1068fdb734a",
    "leased_until_ms": 1790633929437,
    "records": [
      {
        "timestamp_ms": 1790633629168,
        "content": "eyJldmVudCI6InBhZ2VfdmlldyJ9",
        "headers": {
          "content-type": "application/json"
        }
      },
      ...
    ]
  },
  "success": true,
  "errors": [],
  "messages": []
}

A call to consume returns a lease on that batch of events lasting 5 minutes. The client can then take one of three actions:

  • ack the batch, marking it processed so it is never redelivered
  • nack it, signalling that processing failed and the batch should be redelivered
  • extend the lease when more time is needed to finish processing
$ curl -X POST "https://d78b09ee1f50430e9ec92a8af92b0231.k2.cloudflarestorage.com/subscriptions/ee13f761783d3823a447a47b572ebf76/batches/b7e4c9210a3f468d95c2e1068fdb734a/ack" \
  -H "Authorization: Bearer ${CLOUDFLARE_API_TOKEN}" \
  -H "Content-Type: application/json" \
  --data '{ "worker_id": "analytics-1" }'

That covers one consumption model, where work is divided among consumers so each receives a portion of the data. A second option is a separate subscription per consumer — the pub/sub pattern — in which every consumer sees every message. The two can be combined with multiple independent consumer pools. Full API details are in the K2 docs.

Beta limits and pricing

K2 is in public beta for accounts on Workers Paid subscriptions, within the following limits:

  • 10GB maximum storage
  • 30 MB/s produce per stream

Higher limits can be requested through the team's Discord or the limit increase form. No billing applies during the beta. The anticipated pricing once billing starts is:

Pricing

Data Produced

$0.04 / GB

Data Consumed

$0.04 / GB

Data Retained

$0.02 / GB / month

Roadmap

Planned work over the coming months includes:

  • Higher write parallelism, up to multi-GB/s streams
  • Message keys and key-based ordering guarantees
  • Push-based worker consumers
  • An express tier with lower produce and end-to-end latencies
  • Drop-in support for Apache Kafka clients

Feedback is welcome on the Cloudflare Discord.