Semantic search has been the most requested item on Kent C. Dodds' site since the 2021 relaunch, and the request predates that rebuild: the older site carried it as issue #107, opened in 2019. The current tracker held it as issue #5 until recently. It's now live — open the search bar with /, try a natural-language question such as "How did Kent get his first job?" or "What's the best way to learn React?", and press ? for the keyboard shortcuts.

The stack is Cloudflare Workers AI with AI Gateway alongside Vectorize, chosen as a vector database for indexing. Two code paths carry the feature, trimmed here to the essential parts, with GitHub Actions running indexing against content that includes YouTube videos, blog posts, pages and podcasts.

Indexing pipeline

Chunking and content hashing

A stable chunk ID plus a per-chunk payload hash is what makes incremental reindexing practical.

// `source` is the full text content for one document (a single post/page/etc).
const chunkBodies = chunkTextRaw(source, {
	targetChars: 2500,
	overlapChars: 250,
	maxChunkChars: 3500,
})

for (let i = 0; i < chunkBodies.length; i++) {
	const chunkBody = chunkBodies[i] ?? ''
	const vectorId = `${docId}:chunk:${i}`
	const text = `Title: ${title}\nURL: ${url}\n\n${chunkBody}`
	const hash = sha256(text)

	nextManifestChunks.push({ id: vectorId, hash })
}

Here source is a plain string rather than an object — typically the full file contents for MDX-backed documents. YAML-backed sections and transcript-based content get converted into a synthetic text document first, which then serves as source.

chunkTextRaw breaks the string into overlapping chunks, aiming near targetChars while never exceeding maxChunkChars. The overlap keeps context from being severed at a boundary, so information that straddles two chunks still retains its surrounding words. Chunking matters twice over: embedding models impose input limits, and retrieval precision improves when a match points at a relevant passage instead of a whole document. Identifiers of the form docId:chunk:i make updates and deletions deterministic, while the hash allows unchanged chunks to be skipped entirely — no re-embedding, less cost, faster indexing.

Skipping work via the manifest

Chunks whose hash is unchanged never reach the embedding step.

const oldHashesById = new Map(
	(oldManifestDoc?.chunks ?? []).map((c) => [c.id, c.hash]),
)

if (oldHashesById.get(vectorId) === hash) continue

toEmbed.push({
	id: vectorId,
	text,
	metadata: { title, url, snippet: makeSnippet(chunkBody), chunkIndex: i },
})

Embedding and upserting only changed chunks

Changed chunks are embedded and pushed to Vectorize, after which nextManifestChunks is written back to the manifest.

if (toEmbed.length) {
	// getEmbeddings: Array<text> -> Array<vector> (one embedding vector per input text)
	// https://gateway.ai.cloudflare.com/v1/{account_id}/{gateway_id}/workers-ai/{model}
	const vectors = await getEmbeddings({ texts: toEmbed.map((x) => x.text) })

	// https://api.cloudflare.com/client/v4/accounts/{account_id}/vectorize/v2/indexes/{index_name}/upsert
	await vectorizeUpsert({
		// vectorizeUpsert writes or replaces vectors in the index by `id`.
		vectors: toEmbed.map((item, i) => ({
			id: item.id,
			values: vectors[i],
			metadata: item.metadata,
		})),
	})
}

getEmbeddings goes to Workers AI (via AI Gateway in the full implementation) and returns dense numeric vectors per text input. vectorizeUpsert hands those vectors to Vectorize: insert when the vector ID is new, update when it already exists.

Query path

Embedding the query, then overfetching matches

Chunk-level matches are deliberately overfetched at safeTopK * 5, capped at 20, since a single document frequently occupies several of the top slots.

// `K` means "how many nearest neighbors/results to return".
const safeTopK = Math.max(1, Math.min(15, Math.floor(topK)))
const rawTopK = Math.min(15, safeTopK * 5)

const [queryVector] = await getEmbeddings({ texts: [query] })

// https://api.cloudflare.com/client/v4/accounts/{account_id}/vectorize/v2/indexes/{index_name}/query
const { matches } = await queryVectorize({
	vector: queryVector!,
	topK: rawTopK,
	returnMetadata: 'all',
})

queryVectorize performs the vector search itself: the query embedding goes to the Vectorize API, which returns nearest neighbors, each carrying a similarity score and metadata.

Collapsing chunks into documents

Duplicate documents across results are resolved by building a canonical doc ID and retaining only the highest-scoring chunk per document.

const byDocId = new Map<string, SearchResult>()

for (const m of matches) {
	const type = typeof m.metadata?.type === 'string' ? m.metadata.type : 'doc'
	const slug = typeof m.metadata?.slug === 'string' ? m.metadata.slug : m.id
	const canonicalId = `${type}:${slug}`

	const existing = byDocId.get(canonicalId)
	if (!existing || m.score > existing.score) {
		byDocId.set(canonicalId, {
			id: canonicalId,
			score: m.score,
			title: m.metadata?.title as string | undefined,
			url: m.metadata?.url as string | undefined,
			snippet: m.metadata?.snippet as string | undefined,
		})
	}
}

return [...byDocId.values()]
	.sort((a, b) => b.score - a.score)
	.slice(0, safeTopK)