From preview to GA
AI Search is built on Workers AI, Vectorize, R2, and Browser Run, forming a managed index and retrieval pipeline. Cloudflare has used it to run search on its own blog and developer documentation, and developers have applied it to internal documentation and site search since its introduction over a year ago. General availability brings expanded multimodal handling: native image embeddings, OCR for PDFs, and higher file size limits.
Billing for AI Search begins November 1, 2026, with a free tier retained on all Workers plans. A reminder email will go out before billing is enabled. On the free monthly allotment, the previously shared pool of 2,000 queries is replaced by 1,000 semantic queries and 1,000 full-text queries.
Native image retrieval alongside captions
The earlier approach to images was indirect: run object detection, produce a caption, then embed the caption text. Images became searchable only to the extent the caption captured the relevant detail, while texture, screenshot state, chart relationships, layout, and fine visual detail were lost. AI Search now embeds image pixels directly and retains captions for textual understanding.
Native multimodal retrieval is available with the Qwen3-VL-Embedding model. The implementation uses Matryoshka Representation Learning (MRL) so smaller embeddings remain useful while storage stays manageable and search stays fast.
Query behavior depends on the instance's embedding model. If it supports images, a query image is embedded directly and lands in the same vector space as indexed images and text. If the model is text-only, AI Search converts the query image to text with ToMarkdown and searches with the resulting caption. Text-only models therefore get basic multimodal support, while models with native image support carry the full visual signal.

Caption-based understanding | A small bird with black-and-white markings perched among golden fruit and green leaves |
Native Image Retrieval | Can match details omitted from the caption: the geometry of the bird’s white eyebrow stripe, its yellow-green plumage, the mixture of smooth and weathered fruit, the leaves’ deeply ribbed texture, and the image’s warm palette and shallow-focus composition. |
Queries the caption can't express
Captions impose a ceiling on visual search. Anything not anticipated by the caption author is discarded, and specialized or lengthy captions are often needed to cover the likely queries. Direct image embeddings keep the visual characteristics without requiring that guesswork, which makes certain queries practical: describing an image you want to find, supplying an image to find visually similar results, or combining both, as in "a bird with similar markings" or "a bird perched on a leafy branch with plums." Target cases include product discovery, screenshot matching, charts, diagrams, scanned documents, and collections where color, texture, composition, or spatial relationships carry meaning.

Query flow
A query is optionally rewritten, then embedded directly by a multimodal model or captioned first by a text-only model. Vector and keyword search execute in parallel; results are fused and optionally reranked. The top chunks are returned to the caller, or handed to a generation model to produce an answer.
Larger files and scanned PDFs
Text files (Markdown, HTML, CSV, JSON and similar) and PDFs are now accepted up to 10 MiB, up from 4 MiB. Scanned PDFs that contain no extractable text can be handled by enabling OCR, which reads text from each page before chunking and embedding. OCR is available to every account and is billed as image processing ingestion tokens under AI Search pricing.
What ingestion, storage, and queries cost
The pricing model was announced during Agents Week in August 2026 and takes effect with GA. It covers three billable dimensions — content ingested, data stored, and queries run. Parsing, chunking, embedding with Workers AI models, keyword indexing, and reranking are included. There are no instance hours, capacity units, or monthly minimums, and small projects fit within the free monthly allotment. Estimated cost therefore reduces to how much content you index, how much you store, and how many queries you expect.
- Ingestion: one rate per token with the chosen Workers AI embedding model; tokens are counted identically across models. Moving from a text-only model to a multimodal one does not change ingestion cost unless images are also processed, which carries an add-on fee.
- Storage: priced on the size of data in your indices.
- Queries: priced by query type (semantic vs. full-text) and volume.
Pricing | Free monthly allotment (all Workers plans) | |
Ingestion | ||
Base Ingestion | $0.75 / 1M tokens | 5M tokens † |
Image processing (add-on) | +$0.50 / 1M tokens | 5M tokens † |
Storage | ||
Stored data | $2.00 / GB-month | 10 GB |
Query | ||
Semantic (hybrid and vector search) | $0.75 / 1k queries | 1,000 queries |
Full-text | $0.10 / 1k queries | 1,000 queries |
Embedding and Reranking | ||
Ingestion and query | Free with select Workers AI models; third-party billed separately | N/A |
† A single pool of 5M ingestion tokens per month, covering any file type currently supported (e.g., text, images).
On the roadmap
- An ingestion pipeline for full video and audio processing, so customers can search rich media assets.
- A refactor of the keyword search engine to scale better with large data stores, where the current implementation has limits.
- Simpler ways to enable AI Search and create indexes for sites already running on Cloudflare, so AI agents can discover, explore, and consume content more easily.



