From preview to GA

AI Search is built on Workers AI, Vectorize, R2, and Browser Run, forming a managed index and retrieval pipeline. Cloudflare has used it to run search on its own blog and developer documentation, and developers have applied it to internal documentation and site search since its introduction over a year ago. General availability brings expanded multimodal handling: native image embeddings, OCR for PDFs, and higher file size limits.

Billing for AI Search begins November 1, 2026, with a free tier retained on all Workers plans. A reminder email will go out before billing is enabled. On the free monthly allotment, the previously shared pool of 2,000 queries is replaced by 1,000 semantic queries and 1,000 full-text queries.

Native image retrieval alongside captions

The earlier approach to images was indirect: run object detection, produce a caption, then embed the caption text. Images became searchable only to the extent the caption captured the relevant detail, while texture, screenshot state, chart relationships, layout, and fine visual detail were lost. AI Search now embeds image pixels directly and retains captions for textual understanding.

Native multimodal retrieval is available with the Qwen3-VL-Embedding model. The implementation uses Matryoshka Representation Learning (MRL) so smaller embeddings remain useful while storage stays manageable and search stays fast.

Query behavior depends on the instance's embedding model. If it supports images, a query image is embedded directly and lands in the same vector space as indexed images and text. If the model is text-only, AI Search converts the query image to text with ToMarkdown and searches with the resulting caption. Text-only models therefore get basic multimodal support, while models with native image support carry the full visual signal.

Image by joelfotos, via Wikimedia Commons, licensed under CC0 1.0.

Caption-based understanding

A small bird with black-and-white markings perched among golden fruit and green leaves

Native Image Retrieval

Can match details omitted from the caption: the geometry of the bird’s white eyebrow stripe, its yellow-green plumage, the mixture of smooth and weathered fruit, the leaves’ deeply ribbed texture, and the image’s warm palette and shallow-focus composition.

Queries the caption can't express

Captions impose a ceiling on visual search. Anything not anticipated by the caption author is discarded, and specialized or lengthy captions are often needed to cover the likely queries. Direct image embeddings keep the visual characteristics without requiring that guesswork, which makes certain queries practical: describing an image you want to find, supplying an image to find visually similar results, or combining both, as in "a bird with similar markings" or "a bird perched on a leafy branch with plums." Target cases include product discovery, screenshot matching, charts, diagrams, scanned documents, and collections where color, texture, composition, or spatial relationships carry meaning.

The life of an AI Search query

Query flow

A query is optionally rewritten, then embedded directly by a multimodal model or captioned first by a text-only model. Vector and keyword search execute in parallel; results are fused and optionally reranked. The top chunks are returned to the caller, or handed to a generation model to produce an answer.

Larger files and scanned PDFs

Text files (Markdown, HTML, CSV, JSON and similar) and PDFs are now accepted up to 10 MiB, up from 4 MiB. Scanned PDFs that contain no extractable text can be handled by enabling OCR, which reads text from each page before chunking and embedding. OCR is available to every account and is billed as image processing ingestion tokens under AI Search pricing.

What ingestion, storage, and queries cost

The pricing model was announced during Agents Week in August 2026 and takes effect with GA. It covers three billable dimensions — content ingested, data stored, and queries run. Parsing, chunking, embedding with Workers AI models, keyword indexing, and reranking are included. There are no instance hours, capacity units, or monthly minimums, and small projects fit within the free monthly allotment. Estimated cost therefore reduces to how much content you index, how much you store, and how many queries you expect.

  • Ingestion: one rate per token with the chosen Workers AI embedding model; tokens are counted identically across models. Moving from a text-only model to a multimodal one does not change ingestion cost unless images are also processed, which carries an add-on fee.
  • Storage: priced on the size of data in your indices.
  • Queries: priced by query type (semantic vs. full-text) and volume.

Pricing

Free monthly allotment (all Workers plans)

Ingestion

Base Ingestion

$0.75 / 1M tokens

5M tokens †

Image processing (add-on)

+$0.50 / 1M tokens

5M tokens †

Storage

Stored data

$2.00 / GB-month

10 GB

Query

Semantic (hybrid and vector search)

$0.75 / 1k queries

1,000 queries

Full-text

$0.10 / 1k queries

1,000 queries

Embedding and Reranking

Ingestion and query

Free with select Workers AI models; third-party billed separately

N/A

† A single pool of 5M ingestion tokens per month, covering any file type currently supported (e.g., text, images).

On the roadmap

  • An ingestion pipeline for full video and audio processing, so customers can search rich media assets.
  • A refactor of the keyword search engine to scale better with large data stores, where the current implementation has limits.
  • Simpler ways to enable AI Search and create indexes for sites already running on Cloudflare, so AI agents can discover, explore, and consume content more easily.