During Birthday Week 2025 we announced the Cloudflare Data Platform, a suite of products for ingesting, storing and querying analytical data. The platform is now generally available, and it has a new name: Cloudflare Basin.

Basin is a serverless data analytics platform built on Apache Iceberg, the open standard for data lakes, and R2 Object Storage. Its three products cover the analytical lifecycle from ingest to query:

  • Basin Pipelines (formerly Cloudflare Pipelines) receives events from Workers, HTTP or Cloudflare Logpush, transforms them with SQL, and writes them as Apache Iceberg tables or files in R2.
  • Basin Catalog (formerly R2 Data Catalog) manages Iceberg metadata and automatically maintains tables to keep them fast and cost-efficient.
  • Basin SQL (formerly R2 SQL) is the serverless, distributed SQL engine that queries Apache Iceberg tables directly on Cloudflare.

Two developments drove the decision to build the platform: Apache Iceberg's emergence as the standard open table format, which makes data portable across nearly every major query engine, and the migration of analytics data to R2, where the absence of egress charges makes it practical to reach the same data from different tools, teams, regions and cloud providers.

Early adopters included Cloudflare's own billing and infrastructure teams, who used the services for long-term storage and reporting of billing metrics and for ingesting and querying telemetry to measure infrastructure utilisation and efficiency. Other users applied real-time data to e-commerce optimisation.

"We moved our entire company's data pipeline to Basin Pipelines, Catalog, and SQL, replacing a complex AWS S3 and Athena setup with a cleaner, serverless architecture that reliably handles all of our event data. -Dax Raad, Co-Founder, Anomaly

What the year of beta changed

Work since the beta concentrated on three properties that distinguish analytics on the Developer Platform: speed, openness and cost efficiency.

On speed, a Catalog can be created, a Pipeline set up to ingest data, and that data queried with Basin SQL in seconds — a short loop that matters as more data applications are generated by coding agents and would otherwise poll for resources. As datasets grow, Catalog compacts metadata and data files to cut I/O and generates statistics for query planning, while Basin SQL uses those statistics to split queries into smaller tasks distributed across Workers.

On openness, data written through Basin can be read and written with any Iceberg-compatible engine, including PyIceberg, DuckDB, Snowflake and Apache Spark. That portability depends on free egress, which lets developers reach data held in Cloudflare from tools anywhere, in any region or cloud.

"Bobsled is a data product platform that the world's most advanced data teams use to build and distribute AI-ready data to partners, vendors and customers," said Julien Grobbelaar, Head of Platform at Bobsled. "Basin allows us to build data products that can be made accessible in any region of every major data and AI platform, all at production-grade reliability and a fraction of the cost thanks to zero egress fees."

On cost, the serverless architecture supports usage-based pricing: billing applies when Basin ingests, processes or queries data, with no hourly charges or separate infrastructure costs.

Why the name

Cloudflare Data Platform served the first year of availability, but the GA launch called for a name that covers the platform's ambitions, ties the products together and is catchier. A basin collects rivers from many sources at a single point, which maps onto how the products work together: Pipelines brings data into Catalog while Basin SQL makes it queryable. Roughly 20% of Earth's land drains into endorheic basins, a figure that rhymes with the share of the web behind Cloudflare's network.

Lake Vänern, Sweden - A basin created from tectonic activity 200-300 million years ago.

Basin will expand over time with more products managing the rest of the analytical data lifecycle.

Ingestion with Basin Pipelines

Events must be ingested, structured to a schema and written to object storage before they can be queried. Basin Pipelines accepts events through HTTP endpoints or Workers bindings, processes them according to a SQL query, and delivers them to Basin Catalog as Apache Iceberg tables or to R2 as JSON or Parquet files. Users have created tens of thousands of Pipelines since the beta; a common pattern is transforming Cloudflare HTTP logs before storing them.

INSERT INTO http_logs_sink
SELECT
  EdgeResponseStatus,
  to_timestamp_micros(EdgeStartTimestamp) AS event_time,
  upper(ClientRequestMethod) AS method,
  sha256(ClientIP) AS hashed_ip
FROM http_logs_stream
WHERE EdgeResponseStatus >= 400;

Transforming during ingestion shrinks the storage footprint, reduces noise from dynamic sources and helps keep sensitive or unnecessary values out of storage.

Scalability has changed considerably: Pipelines now supports ingesting up to 3GB/s per stream. Integrations with other Cloudflare systems have also widened:

  • Logpush integration: Cloudflare logs can be transformed with SQL and stored as compressed Parquet files or Iceberg tables, ready for Basin SQL or another engine.
  • Schema-aware Worker bindings: wrangler types generates TypeScript types from a stream's schema, catching missing fields and type mismatches before deployment.
  • Visible data-quality errors: the dashboard and GraphQL API surface dropped events and distinguish missing fields, type mismatches, parse failures and null values.
  • Infrastructure as code: Terraform resources cover the catalog, stream, sink and the SQL connecting them.

Planned additions include custom partitioning when writing to Basin Catalog, schema migrations with updatable configuration and Pipelines SQL, support for Iceberg V3 including the Variant type for semi-structured data, and stateful processing for streaming aggregations, joins and incrementally updated materialised views.

Table maintenance in Basin Catalog

Basin Catalog, the first product in the family, is used by thousands of developers for cases ranging from giving DuckDB a structured route to analytics data in R2 to full enterprise data-sharing platforms that lean on zero egress fees and simple APIs. It provides a fully managed Apache Iceberg REST catalog, created in a single step, that performs routine maintenance automatically.

npx wrangler basin catalog create CATALOG_NAME

Automatic compaction was already present when the Data Platform was announced. It has since grown into a wider maintenance set:

  • Per-table compaction policies choose target file sizes based on each table's access pattern.
  • Automatic snapshot expiration removes old Iceberg snapshots under a retention policy while preserving a minimum number of recent snapshots.
  • Unreferenced data-file cleanup reclaims storage as snapshots expire, without a separate Spark maintenance job.
  • Manifest optimisation consolidates and clusters fragmented manifests by partition before compaction, reducing metadata I/O during query planning.

In progress for Catalog: a new compaction approach that efficiently sorts and clusters data for better query performance, more granular auth controls for namespaces and tables, and jurisdiction support for data sovereignty and compliance requirements.

Querying with Basin SQL

Basin SQL is the serverless, distributed query engine for Apache Iceberg tables held in Basin Catalog, built for reading large datasets and scaling automatically across Cloudflare's global network. There are no clusters or resources to provision — only an API that developers and agents can query immediately. At beta it excelled at filtering and exploring large event and time-series tables; it now supports hundreds of functions, among them:

  • Standard and approximate aggregations, GROUP BY, HAVING, and schema-discovery commands
  • More than 190 scalar and aggregate functions across strings, timestamps, regular expressions, cryptography, statistics, arrays, maps and structs
  • CASE expressions, common table expressions, casting, arithmetic and EXPLAIN
  • Inner, outer, semi and anti joins; subqueries; self-joins; multi-table queries
  • DISTINCT, UNION, INTERSECT and EXCEPT
  • Window functions, QUALIFY, grouping sets, rollups and cubes
  • A suite of JSON functions

If one Pipeline delivers application events into a table and another delivers account data, the two can be joined, activity aggregated by customer, the results ranked with a window function and the ranking filtered in a single query.

WITH account_activity AS (
  SELECT
    a.plan,
    e.account_id,
    count(*) AS events,
    approx_distinct(e.user_id) AS active_users
  FROM analytics.events e
  JOIN analytics.accounts a
    ON e.account_id = a.account_id
  WHERE e.event_time >= '2026-09-01T00:00:00Z'
  GROUP BY a.plan, e.account_id
)
SELECT
  *,
  rank() OVER (PARTITION BY plan ORDER BY events DESC) AS activity_rank
FROM account_activity
QUALIFY rank() OVER (PARTITION BY plan ORDER BY events DESC) <= 10;

Queries run from Wrangler or the API, or from an editor built into the Cloudflare dashboard with syntax highlighting and autocomplete, a browser for namespaces and tables, query statistics and plans, and exportable results.

Still to come for Basin SQL: advanced statistics and adaptive scheduling for more efficient query execution, full data definition language support directly from the engine, and Iceberg V3 support including the VARIANT and geospatial types.

Roadmap and platform direction

The team's stated goal is to abstract data infrastructure away entirely, treating storage formats, products, and resources as implementation details. The intended endpoint is a platform where developers start from questions rather than CREATE statements or CLI commands.

Work underway toward that vision includes:

  • Support for the latest Apache Iceberg spec across the entire platform
  • Push-button ingestion sources and destinations, with zero-configuration connections across Cloudflare's developer and observability products
  • Advanced adaptive table-maintenance strategies in Basin Catalog that organize data around real query patterns
  • Continued expansion of SQL compatibility, performance, and observability for increasingly complex analytical workloads
  • More ways to continuously process data in real time and trigger actions based on signals in the data
  • Tools for data compliance and sovereignty rules across the platform

The stated commitment is to keep building on open standards: open formats and protocols, contributions back to upstream projects, and data that stays available to the broader ecosystem.

Availability and migration

Basin Pipelines, Basin Catalog, and Basin SQL are generally available today. They can be used together as an end-to-end platform, or adopted individually to fit an existing architecture. Existing Cloudflare Pipelines, R2 Data Catalog, and R2 SQL configurations continue to work.

The getting started tutorial walks through ingesting events, creating an Apache Iceberg table in Basin Catalog, and querying it with Basin SQL. Product guides, pricing, limits, and integrations are in the Basin documentation.

Feedback is collected in the Cloudflare Developer Discord.