Design Retrieval architecture for a coding copilot for a large engineering org
Junior architecture interview question on Retrieval within RAG.
Read full explanationMetadata filters can silently exclude the chunk that answers the question. Pre-filter vs post-filter, ACL traps, over-filtering, and interview debugging tips.

When a RAG demo fails in an interview, candidates often blame the embedding model or the LLM. In production systems the more common silent failure is a metadata filter that removes the only chunk that could answer the question. Interviewers like this topic because it separates people who have drawn a retrieve-then-generate diagram from people who have shipped multi-tenant retrieval with ACLs, document types, and freshness constraints. This guide walks through how metadata filters hide the right chunk, how pre-filter and post-filter interact with ANN indexes like HNSW, and what strong answers sound like in a senior AI interview.
Start with a concrete failure mode. A support agent asks about refund policy for enterprise customers in the EU. The correct paragraph lives in a PDF tagged region equals EU, product equals enterprise, and doc_type equals policy. Your query embedding is fine. Your top-k similarity search would have ranked that chunk first. Then a filter requires region equals US because an upstream router misclassified the user, or a default tenant filter from a previous session leaked into the request. The right chunk never enters the context window. The model answers from adjacent US policy text and looks confidently wrong. From the outside this looks like a hallucination. From the inside it is filtered retrieval.
Metadata filters exist for good reasons. Tenant isolation, language, document type, product line, ACL groups, time ranges, and PII classification all reduce the searchable space. Without them you either leak private documents across tenants or drown the reranker in irrelevant neighbors. The interview trap is treating filters as a free correctness win. Every filter is a recall trade-off. Interviewers want you to say that out loud and then explain how you measure the trade-off with retrieval evals sliced by filter combinations, not only by global Recall at k.
Understand where the filter runs relative to the vector index. Approximate nearest neighbor structures such as HNSW build a graph over the full embedding space. A pre-filter applies the metadata predicate before or during graph traversal so candidates outside the filter never compete. A post-filter retrieves a larger candidate set first, then drops vectors that fail the predicate. Pre-filtering can be more efficient when selectivity is high and the index supports filtered search well, but aggressive pre-filters on sparse attributes can leave too few viable neighbors and degrade ANN quality. Post-filtering preserves neighborhood quality in the unfiltered space, but if you retrieve top-k equals 20 and 19 are filtered out, you may return almost nothing. Strong candidates describe raising overfetch, or using a hybrid of pre-filter on cheap high-cardinality fields and post-filter on rarer predicates, and they mention that vector database behavior differs across pgvector, Pinecone, Weaviate, Qdrant, and Milvus for filtered HNSW.
Over-filtering is the classic hide-the-chunk bug. Teams stack filters until the intersection is empty for many real queries. Examples include requiring doc_type equals faq when the answer lives only in a runbook, requiring status equals published when the latest correction is still draft, or requiring language equals en when the authoritative policy is bilingual and stored under language equals multi. Another pattern is AND-ing ACL groups incorrectly so a user who belongs to group A or group B must match both. Interview talking point: always log filter predicates beside query text, retrieved ids, and residual candidate counts after filtering. If residual count is zero or near zero, treat it as a first-class incident signal, not a soft miss.
Missing or inconsistent metadata is the twin failure. Chunks without a region field fall out of every region filter. Chunks re-ingested after a schema change may lack new required fields. Partial backfills leave half the corpus searchable under the new filter and half invisible. Parent-child chunking makes this worse when metadata is attached only to the parent document and child spans inherit nothing, or inherit stale tags. A good interview answer describes ingestion contracts: required fields validated at write time, default-deny versus default-allow policy for missing keys, and reindex jobs when the filter schema changes. Default-deny is safer for ACL but more dangerous for recall. Default-allow is the opposite. Say which you choose for security filters versus topical filters.
ACL and multi-tenant filters deserve their own depth. Security filters must be correct even when recall suffers. You never post-filter ACLs after the LLM has already seen private text in a shared cache or in tool traces. Prefer enforcement as early as the vector query, and again at the application layer as defense in depth. For tenant filters, never trust client-supplied tenant ids alone; bind them from the authenticated session. Interviewers often ask how you prevent a prompt-injected document from instructing the agent to drop the tenant filter. Your answer should include server-side filter construction, not filter strings assembled from model output.
Pre-filter versus post-filter is a frequent whiteboard question. Structure the answer. Pre-filter: apply predicate in the index path, reduces candidates early, can hurt recall or ANN quality when predicates are very selective or poorly indexed, good for mandatory tenant and ACL constraints when the store supports filtered search. Post-filter: retrieve more neighbors first, then drop, preserves geometric neighbors, risks empty results unless you overfetch, can waste QPS and latency. Mention that some systems support payload indexes or sparse-aware filtered ANN so pre-filter is closer to exact. Mention hybrid search: BM25 or sparse lexical hits often need the same metadata constraints, and inconsistent filters between dense and sparse channels create another silent miss where one channel finds the chunk and the merge step discards it.
Debugging workflow that impresses interviewers. Reproduce with the exact user query, session ACL, and filter map. Run an unfiltered retrieval and inspect whether the gold chunk appears in the top-n. If yes, binary-search which predicate removes it. Check metadata on that chunk for nulls, wrong enums, timezone-shifted dates, and inheritance bugs. Compare overfetch sizes and residual counts. Add an eval slice for that filter combination. Only then consider embedding or chunking changes. This order matters: filter bugs are cheaper to fix than re-embedding a corpus.
What not to overclaim. Do not claim that metadata filters always improve precision without measuring recall. Do not claim HNSW filtered search is identical across vendors. Do not claim post-filtering is always safe for ACLs if any untrusted component sees pre-filter candidates. Do not claim that tagging every chunk with twenty facets is free; facet cardinality increases ingestion complexity and empty-intersection risk. Do not pretend a reranker can recover a chunk that never entered the candidate set.
Strong closing talking points for the interview. Filters are a recall dial, not a free lunch. Log predicates and residual counts. Separate security filters from topical filters in policy and in code paths. Choose pre-filter versus post-filter with selectivity, index support, and latency budget in mind. Keep metadata schemas versioned with ingestion validation. Build retrieval evals that include filtered scenarios and empty-result rates. If you can tell that story with a production incident where a wrong region filter hid the policy chunk, you will sound like someone who has operated RAG, not someone who only read a tutorial.
Deep explanations with architecture diagrams for every question below.
Junior architecture interview question on Retrieval within RAG.
Read full explanationJunior trade-off interview question on Retrieval within RAG.
Read full explanationMid-Level debugging interview question on Multi-Query Retrieval within RAG.
Read full explanationMid-Level architecture interview question on Hybrid Search within RAG.
Read full explanationMid-Level architecture interview question on Parent Document Retrieval within RAG.
Read full explanationMid-Level evaluation interview question on Query Rewriting within RAG.
Read full explanationMid-Level conceptual interview question on HyDE within RAG.
Read full explanationMid-Level implementation interview question on Parent Document Retrieval within RAG.
Read full explanation