Skip to main content
AI Interview Question
INTERVIEW GUIDEAI News8 questions5 min readOct 4, 2026

Cloudflare AI Search GA: multimodal embeddings, OCR, and hybrid RAG in interviews

Oct 1, 2026: Cloudflare AI Search goes GA with native multimodal embeddings, PDF OCR, hybrid search by default, and billing from Nov 1—interview notes on the caption vs pixel fork.

Cloudflare AI Search GA: multimodal embeddings, OCR, and hybrid RAG in interviews

On October 1, 2026, Cloudflare announced that AI Search is generally available: a managed index and retrieval pipeline that combines Workers AI, Vectorize, R2, and Browser Run. For interview prep, treat this as product news about a packaged hybrid RAG stack—not as proof that any one vendor “wins” retrieval. The useful story is what shipped for multimodal documents, scanned PDFs, hybrid search defaults, and how billing separates ingestion, storage, and queries.

The multimodal change is the sharpest technical hook. Cloudflare says the older image path was caption-only: detect objects, generate a caption, then embed that text, so search could only match what the caption captured. With GA, AI Search embeds image pixels directly for visual retrieval while still retaining captions for textual understanding, and it cites Matryoshka Representation Learning (MRL) to keep richer representations storage-efficient. Native image embedding is available when you choose a multimodal embedding model. Official supported-models docs list Workers AI `@cf/qwen/qwen3-vl-embedding-2b` and Google AI Studio `google-ai-studio/gemini-embedding-2` with image support; the default embedding remains `@cf/qwen/qwen3-embedding-0.6b`, which is text-only. At query time, if the instance model supports images, a query image is embedded into the same vector space as indexed images and text; if the model is text-only, AI Search converts the query image with ToMarkdown and searches on the resulting caption. That distinction is interview gold: “basic multimodal” via captions is not the same as native visual vectors.

File and OCR limits matter for enterprise corpora. Text files such as Markdown, HTML, CSV, and JSON, plus PDFs with OCR enabled, can be up to 10 MiB. PDFs without OCR and other supported formats stay at 4 MiB. OCR is available on every account for scanned PDFs and is billed as image-processing ingestion tokens under AI Search pricing. On the retrieval path, Cloudflare describes the query life as optional rewrite, embed, parallel vector and keyword search, fusion with optional reranking, then return chunks or pass them to a generation model. New instances default to hybrid search (semantic plus full-text), with other index methods selectable at create time.

Model configuration has a production gotcha worth saying out loud. Cloudflare’s models docs state that the embedding model can be chosen when you create a new AI Search instance and is only available to be changed at create time—so swapping embedding strategy later means a new instance and a reindex, not a Settings toggle. Generation models can be changed in Settings and overridden per request. Workers AI embedding and reranking calls made by AI Search are included in AI Search pricing and no longer appear on the Workers AI bill or in AI Gateway logs; generation, query rewriting, and external providers still use your account and gateway.

Billing starts November 1, 2026, with a reminder email beforehand. Cloudflare publishes a free monthly allotment of 5 million ingestion tokens, 10 GB-month of storage, 1,000 semantic queries, and 1,000 full-text queries. Beyond that allotment, published rates are $0.75 per million base ingestion tokens, an additional $0.50 per million for image processing, $2.00 per GB-month storage, $0.75 per 1,000 semantic, vector, or hybrid queries, and $0.10 per 1,000 full-text queries. Docs say ingestion tokens are counted on final chunks after parsing and chunking with the cl100k_base tokenizer, with overlap counted in each chunk, and OCR-extracted text counted as image-processing tokens in addition to base ingestion. Attribute these numbers to Cloudflare’s pricing page; do not invent SLOs or customer savings.

In an AI/ML or RAG interview, lead with the architecture choice: managed hybrid retrieval with a multimodal fork you must configure at create time. Contrast caption-only image search with native multimodal embeddings when visual detail (product texture, screenshots, charts, scanned layout) matters. Mention OCR and the 10 MiB OCR path for scanned PDFs, hybrid as the new default, and the cost model that folds Workers AI embed/rerank into AI Search while leaving generation on the gateway. Stay honest: this is Cloudflare’s GA announcement and docs—not an independent quality audit—and it is a different product from Cloudflare’s separate Web Search API for live web grounding via AI Gateway.

https://blog.cloudflare.com/ai-search-ga/ https://developers.cloudflare.com/changelog/post/2026-10-01-ai-search-generally-available/ https://developers.cloudflare.com/ai-search/platform/limits-pricing/ https://developers.cloudflare.com/ai-search/configuration/models/supported-models/ https://developers.cloudflare.com/ai-search/configuration/models/

CloudflareAI SearchRAGMultimodalEmbeddingsHybrid searchOCRVectorize

Questions in this guide

Deep explanations with architecture diagrams for every question below.