Explain Latency with a production scenario
Principal conceptual interview question on Latency within RAG.
Read full explanationBudget a RAG request stage by stage. Generation drives latency, retrieved tokens drive cost. Tune top-k, rerank depth, ANN dials, and caching against recall.

Interviewers asking about RAG latency and cost want a per-stage budget, not "use a faster model". A RAG request is a chain: an optional query rewrite, query embedding, vector or hybrid search, an optional rerank, prompt assembly, then generation. Each stage adds its own time and its own bill. The strong answer measures every stage, names the one that dominates, and cuts there without quietly losing recall.
This is the RAG-specific companion to the LLM Cost & Latency Interview Guide (https://aiinterviewquestion.com/blog/llm-cost-latency-interview-guide), which covers caching, routing, and streaming for LLM calls in general. Index internals are in the Vector Database Interview Guide (https://aiinterviewquestion.com/blog/vector-database-interview-guide).
Start with the request path
Draw the path and put a timer on each box. OpenAI's latency optimization guide uses this shape for its worked example: a customer-service bot that rewrites the last message into a self-contained query, checks whether retrieval is needed, retrieves, then generates the answer. In that example the rewrite and the retrieval check were separate model calls. The changes were to combine them into one request, move that step to a smaller fine-tuned model, and run independent steps in parallel. OpenAI's own conclusion is that the right split depends on testing with production examples. That is the answer shape interviewers want: trace, attribute, then change one stage.
Log per request: rewrite time, embed time, search time, rerank time, input tokens, output tokens, time to first token, and total time. Without per-stage numbers, "RAG is slow" cannot be debugged.
Generation usually dominates latency
OpenAI's guide says generating tokens is almost always the highest latency step when using an LLM. Its heuristic is that cutting 50% of output tokens may cut about 50% of latency. It says the opposite about input: cutting 50% of the prompt may only give a 1 to 5% latency improvement, unless the context is truly massive. So the first latency levers are output length (ask for concise answers, set max tokens, trim structured output) and model size. OpenAI names model size as the main factor in inference speed, and says smaller models usually run faster and cheaper.
Streaming does not make the model generate faster. OpenAI calls it the single most effective way to make users wait less, because it cuts the waiting time to a second or less. Say which number you are optimizing: time to first token is what the user feels, total time is what your capacity plan pays for.
Retrieved context dominates cost
Input tokens may barely move latency, but API models bill them on every request. In a RAG prompt most of those tokens are retrieved chunks, so top-k times chunk size is a line on the bill.
More context is also not free quality. Liu et al., Lost in the Middle (TACL 2024), ran a retriever-reader setup on NaturalQuestions-Open with Contriever retrieval. Reader performance saturated long before retriever recall did. Using 50 retrieved documents instead of 20 improved accuracy by only about 1.5% for GPT-3.5-Turbo and about 1% for claude-1.3, while significantly increasing the input context. The same paper found accuracy is often highest when the relevant passage is at the beginning or end of the context and drops when it sits in the middle. Those are 2023 models on one dataset, so do not quote them as a law. Quote them as the reason you pick k from a measured quality-versus-k curve on your own eval set, not from the largest context window you can afford.
Rerank buys precision, not free speed
A cross-encoder scores every query and document pair. Sentence Transformers says scoring thousands or millions of pairs would be rather slow, which is why its pipeline retrieves about 100 candidates first and reranks only those. Reranker cost grows with the candidate count. In Anthropic's Contextual Retrieval experiments, the top 150 retrieved chunks were reranked down to the top 20. Anthropic says reranking inevitably adds a small amount of latency, even though the reranker scores chunks in parallel, and that there is a trade-off between reranking more chunks for quality and fewer for lower latency and cost. Anthropic also argues reranking can lower downstream cost and latency because the model processes less information. Both effects are real, so name the two knobs separately: candidates in, passages out. The full two-stage argument is in When to add a reranker (https://aiinterviewquestion.com/blog/when-to-add-reranker-rag-interview).
Search and index dials are recall dials
pgvector does exact nearest neighbor search by default, which gives perfect recall. Its README says HNSW has a better speed-recall trade-off than IVFFlat but slower builds and more memory, and IVFFlat builds faster and uses less memory. hnsw.ef_search defaults to 40 and ivfflat.probes defaults to 1. A higher probes value gives better recall at the cost of speed. Turning these down makes search faster by searching less, so check recall on the same eval set before and after.
Filtering hides a cost trap. pgvector applies filters after the approximate index scan. Its README example: if a condition matches 10% of rows, HNSW with the default ef_search of 40 returns only about 4 matching rows on average. Iterative index scans, added in 0.8.0, scan more of the index until enough results are found. A filtered query can look fast because it returned too little. That failure is covered in Metadata filters that hide the right chunk (https://aiinterviewquestion.com/blog/metadata-filters-that-hide-the-right-chunk-rag-interview-guide).
Embedding size is the storage and memory line. OpenAI's text-embedding-3-small and text-embedding-3-large default to 1536 and 3072 dimensions, and the dimensions parameter shortens them. OpenAI says larger embeddings generally cost more and use more compute, memory, and storage. It also reports that on MTEB, text-embedding-3-large shortened to 256 dimensions still outperforms an unshortened text-embedding-ada-002 at 1536. Changing dimensions means re-embedding the corpus, as in Picking an embedding model for retrieval (https://aiinterviewquestion.com/blog/picking-embedding-model-retrieval). For that offline work, OpenAI's Batch API supports embeddings at a 50% discount versus synchronous calls, with completion within 24 hours. That suits ingestion and re-embedding, not the live query path.
Caching: what can and cannot be cached
Prompt caching reuses a prompt prefix the model has already processed. OpenAI requires the entire rendered prefix to match, sets the minimum cacheable length at 1,024 tokens for GPT-5.6 and later, and on those models prices cache writes at 1.25 times the uncached input rate and reads at 0.1 times on most of them. Anthropic caches the prefix in the order tools, system, messages, needs 100% identical segments for a hit, and prices 5-minute cache writes at 1.25 times base input and reads at 0.1 times, with per-model exceptions.
Retrieved chunks change with every query, so they usually sit outside the reusable prefix. OpenAI's latency guide gives the RAG version of the rule: put dynamic parts such as RAG results and history later in the prompt, so the shared prefix stays cache-friendly. Put stable instructions, tool definitions, and examples first. The interview point is that caching cuts the cost of the static part of the prompt. It does not make per-query retrieved context free.
For a small corpus, retrieval may not be the cheapest design. Anthropic wrote in 2024 that a knowledge base under 200,000 tokens, about 500 pages, can go into the prompt whole, and claimed prompt caching cut latency by more than 2x and cost by up to 90%. That is a vendor claim. Anthropic itself says a growing knowledge base needs a more scalable solution, and per-tenant access rules make one shared prompt harder, which is where Multi-tenant RAG isolation (https://aiinterviewquestion.com/blog/multi-tenant-rag-isolation) starts.
Make fewer calls
OpenAI notes that every request adds round-trip latency, and recommends folding sequential steps into one request and not defaulting to an LLM when hard-coding, pre-computing, or classical caching will do. In OpenAI's example, the retrieval check returns false for "Thank you!", so that turn skips search entirely, and the rewrite and the check share one call. Every LLM step before retrieval is latency added in series.
Sample interview answer
"First I trace one slow request and one expensive request and split them into rewrite, embed, search, rerank, and generation, with time to first token and total time. If generation dominates, I cut output tokens and stream. OpenAI's heuristic is that halving output roughly halves latency, while halving input gives only 1 to 5%. If cost is the problem, I look at top-k times chunk size, because those tokens are paid on every call, and Lost in the Middle found that going from 20 to 50 documents added only about 1 to 1.5% accuracy on their QA set. I tune k and the rerank candidate count against a quality-versus-k curve on our eval set. I keep the static prompt prefix stable so prompt caching applies, and accept that retrieved chunks stay uncached. On the index, ef_search and probes are recall dials, so I measure recall before and after, and I check filtered queries that return too few rows. Re-embedding runs through a batch API. Every change goes through the same eval gate, so a cheaper pipeline that drops recall does not ship."
The close: measure latency and cost per stage, and check every cut against recall and answer quality, not just the dashboard. If retrieval quality itself is the question, hand off to Dense vs hybrid search (https://aiinterviewquestion.com/blog/dense-vs-hybrid-search) or How to debug a RAG pipeline (https://aiinterviewquestion.com/blog/how-to-debug-a-rag-pipeline-retrieval-ranking-and-grounding).
https://developers.openai.com/api/docs/guides/latency-optimization https://developers.openai.com/api/docs/guides/prompt-caching https://docs.anthropic.com/en/docs/build-with-claude/prompt-caching https://arxiv.org/abs/2307.03172 https://aclanthology.org/2024.tacl-1.9.pdf https://github.com/pgvector/pgvector https://www.sbert.net/examples/sentence_transformer/applications/retrieve_rerank/README.html https://www.anthropic.com/news/contextual-retrieval https://developers.openai.com/api/docs/guides/batch https://developers.openai.com/api/docs/guides/embeddings
Deep explanations with architecture diagrams for every question below.
Principal conceptual interview question on Latency within RAG.
Read full explanationPrincipal production incident interview question on Latency within RAG.
Read full explanationPrincipal system design interview question on Latency within RAG.
Read full explanationSenior trade-off interview question on Context Precision within RAG.
Read full explanationSenior trade-off interview question on Corrective RAG within RAG.
Read full explanationSenior debugging interview question on Corrective RAG within RAG.
Read full explanationSenior debugging interview question on Faithfulness within RAG.
Read full explanationStaff debugging interview question on Indexing Pipelines within RAG.
Read full explanation