Evaluate Query Rewriting quality in an AI search product
Mid-Level evaluation interview question on Query Rewriting within RAG.
Read full explanationA composite story of a failed RAG deep-dive interview: chunking, retrieval checks, reranking, query rewriting, and evaluation, and what to study instead.

A quick note first. The candidate in this story is an illustrative composite, not a real person, and the interview is not a report of any real company's process or result. It is a realistic pattern, told in first person so the mistakes are easy to recognize.
I walked into the RAG deep-dive thinking it would be easy. I had built a document chatbot. It answered questions, it cited sources, and the demo looked good. The interviewer opened by asking me to walk through it, then spent the next forty minutes asking why each piece was the way it was. I had answers for what I built. I did not have answers for why.
The first crack was chunking. The interviewer asked how I split documents. I said fixed-size chunks with some overlap, because that was the default in the tutorial I followed. Why that size? Why that overlap? What happens to a table, or a section heading that ends up in a different chunk from its paragraph? I had never looked. What I would study differently is to treat chunking as a design choice, not a setting: inspect real chunks, see what breaks, and be able to say what you traded. The chunking and overlap guide at /blog/chunking-and-overlap-rag-interview-guide is where I would start now.
The second crack was retrieval quality. I was asked how I knew retrieval was working. I said the answers looked right. The follow-up was what happens when the right passage is not in the top results. I had no way to tell, because I had only ever read the final answers. What I would study differently is checking retrieval on its own, before generation: keep a small set of questions where you know which passage should come back, and look at what actually comes back. The RAG interview guide at /blog/rag-interview-guide covers the full pipeline, and the RAG debugging walkthrough at /blog/how-to-debug-a-rag-pipeline-retrieval-ranking-and-grounding shows how to split a failure into retrieval, ranking, and grounding.
The third crack was ranking. The interviewer asked whether I would add a reranker. I said yes, because rerankers improve results. Then came the harder question: what problem would it solve in my system, and what would it cost? I could not say. What I would study differently is the job of each stage. A retriever decides what is even in the candidate set. A reranker reorders that set. If the right passage never came back, reordering will not save it. The guide at /blog/when-to-add-reranker-rag-interview walks through when that second stage earns its latency.
The fourth crack was the queries themselves. The interviewer typed a vague, two-part question and asked what my system would do with it. It would embed the whole thing and hope. What I would study differently is query handling as its own step: rewriting a vague question, splitting a multi-part one, and knowing when that extra step is worth it. The query rewriting and decomposition guide at /blog/query-rewriting-decomposition-rag-interview-guide is the reading I skipped and should not have.
The last crack was evaluation. Asked how I would decide the system was good enough to ship, I said I would test it with users. That was the answer that ended the round for me. What I would study differently is a simple evaluation story: separate checks for whether retrieval found the right context, whether the answer stayed faithful to it, and whether it answered the question, run on a fixed set every time something changes. The LLM evaluation guide at /blog/llm-evaluation-interview-guide lays this out.
Looking back, my project was not the problem. The problem was that I had built it by following defaults and never broken it on purpose. The prep that fixed it was rebuilding the same project slowly: changing one stage at a time, writing down what moved, and keeping notes on every failure. In the next deep-dive, I still got questions I could not answer. But I could say which stage the question belonged to and what I would measure to find out, and that is the conversation a deep-dive is really trying to have.
Deep explanations with architecture diagrams for every question below.
Mid-Level evaluation interview question on Query Rewriting within RAG.
Read full explanationJunior architecture interview question on Retrieval within RAG.
Read full explanationStaff evaluation interview question on Access Control within RAG.
Read full explanationMid-Level evaluation interview question on Agentic RAG within RAG.
Read full explanationJunior evaluation interview question on Chunking within RAG.
Read full explanationJunior evaluation interview question on Generation within RAG.
Read full explanationSenior evaluation interview question on Groundedness within RAG.
Read full explanationSenior evaluation interview question on Self-RAG within RAG.
Read full explanation