Debug Multi-Query Retrieval regression in a multi-tenant AI support platform
Mid-Level debugging interview question on Multi-Query Retrieval within RAG.
Read full explanationWrong RAG answers are usually retrieval, filters, or ranking, not the model. Split the failure, then fix chunks, top-k, rerank, and citations.

A RAG pipeline fails in layers. The retriever can miss the document. A metadata filter can delete the hit. A similarity cutoff can treat a good chunk as noise. The reranker can prefer a keyword-stuffed neighbor. The model can then write a fluent answer that never uses the context you fetched. Debugging means freezing everything downstream and asking whether this stage saw the right text.
Interviewers score that habit. "I would add a better prompt" is a junior answer. "I would log retrieved ids, scores, and the prompt, then separate retrieval hit rate from groundedness" is a senior one.
Separate retrieval failures from generation failures
Do not start in the chat UI. Take one bad answer and call retrieval alone. Save chunk ids, scores, metadata, and the query after any rewrite. Label the failure before you edit prompts.
Retrieval miss. The gold passage is not in the candidate set. No prompt can cite text that never arrived. Fix the index, the query, or the filters.
Retrieval hit, generation miss. The gold passage is in the prompt and the model still contradicts it, mixes it with a neighbor, or answers from memory. Fix the prompt contract, context order, the citation rule, or the model.
Abstain failure. The corpus does not contain the answer, and the system invented one. That is grounding, not recall.
Paste only the gold chunk into the prompt and ask again. If the model is right in isolation and wrong in the full pipeline, the bug is upstream or in how context is packed. If it is still wrong on the gold chunk, stop looking at the vector database.
Log three things on every request: the rewritten query, the retrieved set (id, score, source, chunk hash), and the final prompt. Without them, a stale index and a hallucinated sentence look the same.
Chunking and embedding mismatch
A common "dumb model" ticket is a chunk that does not hold a complete fact. Split a refund policy so the window is in one chunk and the exceptions are in the next, and retrieval returns whichever half is closer to the query. The generator answers from a fragment.
Read the stored chunk, not the source PDF. Check four mismatches.
Boundaries. Fixed windows cut tables, lists, and conditions. Prefer heading-aware splits, a little overlap, and a parent section id so a hit can expand to its neighbors.
Embedded text versus prompt text. Embedding a title plus summary while showing a different body sends neighbors at the summary and the model at the body. Embed the text you will show, or log both strings.
Model drift. You cannot query an index with a different embedding model, even from the same vendor. A dimension change fails loudly. The quiet failure is the same dimension and a new model: scores look fine and the neighbors are wrong. Pin the model id on the index and refuse a mismatch.
Identifiers. Error codes, SKUs, clause numbers, and API names often lose to dense search on short queries. If those queries fail, add a keyword leg or a query rewrite. A larger top-k will not create a lexical match the embedding never stored.
Re-embed a few failing chunks with the current model. If the new vectors find the gold passage and the old ones do not, reindex. Do not retune the prompt.
Metadata filters that drop the right document
Filters cause the most silent misses. Tenant, `status: published`, language, or a date range is a hard predicate. The gold document can be the nearest neighbor and still vanish because it is `draft`, tagged `en-GB` while the filter says `en`, or owned by another workspace.
Run the same query with filters off. If the gold id appears only then, the predicate is the bug. Then print the metadata stored on that id. A missing field often fails closed and looks like an empty corpus.
Some approximate indexes filter before the graph search and never visit a rare tag. If filtered queries are empty or slow while unfiltered queries look healthy, compare a pre-filter plan with a post-filter plan and report which one you measured.
Top-k and similarity thresholds
Top-k always returns k neighbors, including garbage on an off-topic query. A threshold can return zero, which is correct when the corpus has no answer. One global cutoff is rarely calibrated. Do not memorize "cosine above 0.8." Scores depend on the model, normalization, and whether the engine returns distance or similarity. Compare them only inside one index build.
On real failures, plot scores of human-relevant chunks against the top score of queries that should abstain. Use a threshold only if those distributions separate. If they overlap, a cutoff either hides good chunks or admits distractors. Raise candidate k and rerank instead.
Candidate k and context k are different. Retrieve on the order of 20 to 50 chunks for the reranker. Put a handful in the prompt after reranking, and add neighbor chunks when the fact was split. Dumping the raw list into the model costs more, adds distraction, and hides ranking bugs the model sometimes lucks through.
Reranking
If the gold chunk is in the candidate set but not in the prompt, that is a ranking miss, not a retrieval miss. A bi-encoder scores the query and the document apart. It is weak on contrasts such as "not covered" versus "covered." A cross-encoder reranker reads them together.
For about ten queries, record gold rank before rerank, gold rank after, and whether the chunk entered the prompt. If rank never improves, the chunk text is incomplete or the reranker is scoring a different string than the embedding. If offline gains disappear in production, you are reranking a stale or unfiltered list.
Say the latency bound in an interview: rerank the candidate set, not the corpus, and skip the reranker when the top vector score already stands clear and the citation check passes.
Citation and grounding checks
Fluency is not support. Require chunk ids on claims, and drop any citation that was not retrieved. That removes fabricated sources.
Word overlap is a weak judge. It fails on a good paraphrase and passes when the model copies a sentence and adds "not." For numbers, dates, and names, be strict: every numeral in the answer should appear in a cited chunk unless the user asked for arithmetic on those numbers. A smaller judge model can score the other claims.
On failure, abstain or narrow the answer. The line to memorize is retrieve, cite, verify, or say the corpus does not support an answer.
Eval sets you can debug against
Do not tune top-k from one screenshot. Freeze 30 to 100 queries from production failures and from documents you know are indexed. Store the question, gold chunk ids, the expected fact, and whether the system must abstain.
Keep the metrics apart:
Recall@k on the candidate set, before the prompt cutoff. Prompt hit rate after rerank and trimming. Groundedness: the expected fact is in a cited chunk, and extra claims are not invented. Abstention on out-of-corpus questions.
Change one component and rerun the same set. A new chunker, embedder, and prompt in one release makes the next drop unreadable. Keep the old index so you can separate reindexing from generation.
Slice the set. Average recall can hide a total miss on tables, date filters, or second-hop questions. Those slices are what you describe in a design review.
What to say in an interview
Walk one example left to right: the question, the fact the corpus should prove, and the stage that lost it. Name the logs: rewritten query, ids and scores, filters, rerank order, prompt, citation check.
Then one causal sentence. "Recall@20 was fine and the prompt missed the chunk, so I would rerank, not re-embed." Or: "Ingestion stored `product=payments` and the filter asked for `billing`, so the policy never entered search. I would fix that contract and add an eval item that fails if the filter hides it." Or: "The number was in the chunk and the answer changed it, so I would check numerals and abstain on mismatch."
Close on the trade. Hybrid search and reranking buy recall with latency. Stricter thresholds and citation checks buy fewer confident wrong answers by abstaining more. Say which pain you are reducing, and which frozen metric would move if the change worked.
Deep explanations with architecture diagrams for every question below.
Mid-Level debugging interview question on Multi-Query Retrieval within RAG.
Read full explanationSenior debugging interview question on Faithfulness within RAG.
Read full explanationMid-Level evaluation interview question on Query Rewriting within RAG.
Read full explanationStaff debugging interview question on Indexing Pipelines within RAG.
Read full explanationJunior trade-off interview question on Retrieval within RAG.
Read full explanationSenior debugging interview question on Corrective RAG within RAG.
Read full explanationJunior debugging interview question on PDF Extraction within RAG.
Read full explanationJunior debugging interview question on RAG Architecture within RAG.
Read full explanation