Explain Multi-Tenancy with a production scenario
Senior conceptual interview question on Multi-Tenancy within RAG.
Read full explanationA prompt line is not an ACL. Bound the store before the model sees another tenant's chunks, and judge the retrieve response, not the answer.

Interviewers asking about multi-tenant RAG isolation want the access check on the retrieve call, not in the prompt. A prompt line is not an access control list. OWASP LLM02:2025 says system-prompt restrictions on what the model returns may not always be honored and could be bypassed via prompt injection or other methods. The older LLM06 page makes the same claim.
The leak in this guide is another tenant's chunk coming back because the store was not bounded before the model. That is not the metadata-filter problem where a filter hides the right chunk. That one is recall. This one is cross-tenant leakage. The OWASP RAG cheat sheet, Section 4, describes a classified document chunked into a shared vector store without per-chunk access metadata, then retrieved by an unauthorized query. The exact point is that access control must be enforced before content reaches the model. Section 6 says not to rely solely on post-retrieval filtering, which means retrieve all, then filter. Section 6 also says to audit chunk isolation with cross-tenant test queries and verify zero cross-boundary results. Section 12's CI case is whether tenant A's query returns tenant B's chunks. These sources do not give a leak rate.
Pinecone's vendor documentation is not an independent security lab. That page says records live in namespaces, upserts and queries always target one namespace, and namespaces are stored separately, which it calls physical isolation. The same page contrasts a weaker path: if isolation is not strict, use one namespace plus a metadata filter at query time. That filter path is not the isolation they describe. The Python SDK page says omitting the namespace defaults to the empty string and searches only that default namespace. The read succeeds and returns nothing. Nothing is raised or logged. Upserts without a namespace go into that empty-string namespace. The Query API reference does not list namespace as required. topK is required. Do not smooth that conflict. The SDK page also says query_namespaces() can query each namespace and merge the results. Omitting the namespace does not search every tenant.
Weaviate's vendor docs say each tenant is a separate shard and data in one tenant is not visible to another. On Get, the tenant name is required if multi-tenancy is enabled. These pages do not give the error text for a Get that omits the tenant, so do not invent one. On the write side, by default an insert into a non-existent tenant returns an error. If autoTenantCreation is true, TenantOne, tenantOne, and a typo create three tenants.
Qdrant's vendor docs say is_tenant is a storage layout, not an access control. Partition with a payload field, then filter. A request without that filter scans all groups and is slower. It does not refuse the request. Do not call is_tenant an ACL.
Elastic's document-level security is vendor documentation, not an independent lab. It restricts read access so non-matching documents are never returned. A role with no document query grants all documents. Omitting the query parameter disables document-level security. Document-level security does not apply to write APIs. Each document must have the username or role name associated with it, so that this information can be used by the role query for document level security.
Azure AI Search vendor docs describe two different paths. Do not merge them into one mechanism. Security trimming treats the principal as a string in a filter. There is no authentication through that principal. Document-level authorization is the filter on every query. A separate preview page describes x-ms-query-source-authorization, where the service returns only documents the synced permission metadata allows.
The interview check is reasoning, grounded in the OWASP line that access control must be enforced before content reaches the model. The state to check is the chunk ids and tenant metadata on the retrieve response, not a model sentence that says it stayed in tenant A. The model's answer text is not the pass or fail signal.
https://raw.githubusercontent.com/OWASP/www-project-top-10-for-large-language-model-applications/main/2_0_vulns/LLM02_SensitiveInformationDisclosure.md https://genai.owasp.org/llmrisk2023-24/llm06-sensitive-information-disclosure/ https://cheatsheetseries.owasp.org/cheatsheets/RAG_Security_Cheat_Sheet.html https://docs.pinecone.io/guides/index-data/implement-multitenancy https://sdk.pinecone.io/python/how-to/vectors/namespaces.html https://docs.pinecone.io/reference/api/2026-04/data-plane/query https://docs.weaviate.io/weaviate/manage-collections/multi-tenancy https://docs.weaviate.io/weaviate/api/graphql/get https://qdrant.tech/documentation/manage-data/multitenancy/ https://www.elastic.co/docs/deploy-manage/users-roles/cluster-or-deployment-auth/controlling-access-at-document-field-level https://learn.microsoft.com/en-us/azure/search/search-security-trimming-for-azure-search https://learn.microsoft.com/en-us/azure/search/search-document-level-access-overview
Deep explanations with architecture diagrams for every question below.
Senior conceptual interview question on Multi-Tenancy within RAG.
Read full explanationStaff scenario interview question on Multi-Tenancy within RAG.
Read full explanationSenior system design interview question on Multi-Tenancy within RAG.
Read full explanationJunior architecture interview question on Retrieval within RAG.
Read full explanationStaff evaluation interview question on Access Control within RAG.
Read full explanationStaff implementation interview question on Access Control within RAG.
Read full explanationJunior trade-off interview question on Retrieval within RAG.
Read full explanationStaff security interview question on Access Control within RAG.
Read full explanation