Skip to main content
AI Interview Question
INTERVIEW GUIDERAG8 questions10 min readOct 5, 2026

Query rewriting in RAG: multi-query, HyDE, step-back, decompose

Query rewriting in RAG: contextualize follow-ups, then multi-query with RRF, HyDE, step-back, or sub-questions. Know the drift and latency interviewers probe.

Query rewriting in RAG: multi-query, HyDE, step-back, decompose

Interviewers ask about query rewriting to see whether you treat the user's question as fixed input or as something you can improve before search. A strong answer names the problem each rewrite solves, merges results properly, and states the cost in latency and in drift from what the user meant.

Put simply: the words a user types often do not match the right document. A follow-up like "How long does it cover?" has no subject, and a short question looks nothing like the passage that answers it. Rewriting changes the search input; the answer must still address the original question.

Five techniques follow: contextualization, multi-query with fusion, HyDE, step-back, and decomposition. For hybrid search and fusion background, see Dense vs hybrid search (https://aiinterviewquestion.com/blog/dense-vs-hybrid-search).

Why rewrite the query at all

Ma et al. (Rewrite-Retrieve-Read, 2023) start from the point that "there is inevitably a gap between the input text and the needed knowledge in retrieval", and add a rewrite step before retrieve-then-read. They tried a prompted LLM rewriter and a small T5-large rewriter trained with reinforcement learning (PPO) on the reader's feedback. With Bing as the retriever and ChatGPT and Vicuna-13B as readers, they report that "query rewriting consistently improves the retrieve-augmented LLM performance" and that "the smaller language model can be competent for query rewriting". Their stated limit: adding training makes direct transfer harder than with few-shot prompting.

So a rewriter is a separate component with its own failure modes, and you evaluate it on its own.

Contextualize follow-up questions first

The most common production rewrite turns a follow-up into a standalone search query. OpenAI's latency optimization guide works through a customer-service RAG bot whose query contextualization prompt "re-writes user query to be a self-contained search query". With the history "What is your return policy?" and the follow-up "How long does it cover?", the rewrite is "How long does the return policy cover?". The raw follow-up never mentions returns, so it makes a poor search query.

The guide then makes this cheaper. Because the retrieval check (does this turn need search at all?) needs the contextualized query, it merges both into one prompt that returns JSON with a query and a retrieval flag, "to make fewer requests". It switches that prompt to a smaller, fine-tuned GPT-3.5 "to process tokens faster", and runs the retrieval checks in parallel with the reasoning step.

Multi-query: more variants, then a real merge

Multi-query retrieval asks an LLM for several versions of the question, retrieves for each, and merges the results, since one phrasing can miss what another hits. LangChain's MultiQueryRetriever: "Given a query, use an LLM to write a set of queries. Retrieve docs for each query. Return the unique union of all retrieved docs." Its default prompt asks for 3 versions of the question.

Two details in its current source matter. include_original defaults to False, so the user's own query is not searched. And the unique union only removes duplicates. It does not rank-fuse, so results come back in concatenation order, and you still need fusion or a reranker.

The standard fusion step is Reciprocal Rank Fusion (Cormack, Clarke, and Buettcher, SIGIR 2009). A document's score is the sum, over every list it appears in, of 1 / (k + rank), with k = 60, "fixed during a pilot investigation". It uses only ranks: no training, no score calibration. The authors found RRF "almost invariably improved on the best of the combined results" and equaled or beat Condorcet Fuse and CombMNZ. Then rerank the top candidates against the original question; see When to add a reranker (https://aiinterviewquestion.com/blog/when-to-add-reranker-rag-interview).

RAG-Fusion (Rackauckas, Infineon, 2024) packages this pattern. It is a single-company case study with manual evaluation, not a benchmark. Answers were more accurate and comprehensive, but "some answers strayed off topic when the generated queries' relevance to the original query is insufficient." Over ten back-to-back runs of one query it averaged 34.62 seconds against 19.52 for plain RAG, "1.77 times longer", and the author says exact runtimes are hard to generalize. Suggested fixes: a locally hosted LLM and fewer queries.

No source gives a best number of variants. Say you would tune it on your own eval set, watching recall and drift.

HyDE: search with a fake answer

HyDE (Gao, Ma, Lin, and Callan, 2022) is for zero-shot dense retrieval with no relevance labels, where a short question does not look like the passage that answers it. An instruction-following LLM is told to "write a paragraph that answers the question". That output "is not real, can contain factual errors but is like a relevant document". An unsupervised encoder (Contriever) embeds it, and that vector searches the real corpus. The dense bottleneck acts as "a lossy compressor, where the extra (hallucinated) details are filtered out", the authors argue. They average several samples' embeddings and can include the query itself.

In their setup (InstructGPT plus Contriever), HyDE "significantly outperforms" unsupervised Contriever and is comparable to fine-tuned retrievers across web search, QA, fact verification, and several languages. Limits: it assumes the query is not ambiguous, and with an in-domain fine-tuned encoder, "less powerful instruction LMs can negatively impact the overall performance", though the drops were small.

Query2doc (Wang, Yang, and Wei, 2023) applies a similar idea to BM25: it appends few-shot LLM pseudo-documents to the query and reports gains of 3% to 15% on MS-MARCO and TREC DL without fine-tuning. It warns that "LLM generations may contain factual errors", and notes HyDE "implicitly assumes that the groundtruth document and pseudo-documents express the same semantics in different words, which may not hold for some queries." In LlamaIndex, HyDEQueryTransform defaults to include_original=True, embedding the original query alongside the hypothetical passage.

The interview point: the pseudo-document is a search key, not an answer. It can carry made-up dates, names, or numbers that pull retrieval toward the wrong documents, so never show it to the user. That rule is our advice, based on the paper's warning that the text "is not real". Practice: Explain HyDE with a production scenario (https://aiinterviewquestion.com/questions/explain-hyde-with-a-production-scenario).

Step-back: ask the broader question

Step-back prompting (Zheng et al., Google DeepMind, 2023) moves up a level before retrieving. A step-back question is "a derived question from the original question at a higher level of abstraction." Example: "Estella Leopold went to which school between Aug 1954 and Nov 1954?" becomes "What was Estella Leopold's education history?" The step-back question "is used to retrieve relevant facts" that ground the final reasoning.

With PaLM-2L on TimeQA, they report 41.5% for the baseline, 57.4% with regular RAG, and 68.7% with Step-Back plus RAG. On the Hard split, RAG lifted accuracy only from 40.4% to 46.8%, while Step-Back plus RAG reached 62.3%. These come from their 2023 setup, not a forecast for today's models.

Interviewers like the error analysis. On TimeQA the step-back step "rarely fails". More than half the errors were reasoning errors, and 45% were retrieval errors, a category defined as missing the right information "despite that the step-back question is on target." So a good rewrite does not guarantee good retrieval. On MMLU high-school Physics, Step-Back fixed 20.5% of the baseline's errors but introduced 11.9% new ones.

Decomposition: parallel or sequential sub-questions

Comparisons and multi-hop questions need facts no single chunk holds. Decomposition splits the question into sub-questions, retrieves for each, and combines the answers. The key choice: are the sub-questions independent?

Parallel. LlamaIndex's SubQuestionQueryEngine "breaks down the complex query into sub questions for each relevant data source", then synthesizes the intermediate answers. Its docs example asks 3 sub-questions: Paul Graham's work before, during, and after YC. The notebook's one published trace (2024, illustrative only, not a benchmark) teaches two lessons. One sub-question took about 62 seconds, the others about 1.5 to 2.4, and the whole query about 66: the slowest branch set wall-clock time. And one sub-answer ("Paul Graham worked on writing essays and working on YC before YC") contradicts itself, yet flows straight into the final synthesis.

Sequential. When hop two needs hop one's answer, you cannot plan every sub-question up front. Least-to-most prompting (Zhou et al., 2022) is built to "break down a complex problem into a series of simpler subproblems and then solve them in sequence". Self-ask (Press et al., 2022) has the model ask and answer its own follow-up questions, so a search engine can answer them. It also defines the compositionality gap: how often a model answers every sub-question correctly but gets the overall answer wrong. In the GPT-3 family, that gap did not shrink as models grew. IRCoT (Trivedi et al., 2022) states the core problem: what to retrieve "depends on what has already been derived". Interleaving retrieval with chain-of-thought steps gave GPT-3 gains of up to 21 points in retrieval and 15 in QA on four multi-step QA datasets; one-step retrieve-and-read "is insufficient for multi-step QA".

Sequential decomposition costs more round trips, and (our inference, not a sourced result) a wrong early hop carries into every later hop. If iterative retrieval grows into an agent loop, see Advanced RAG patterns (https://aiinterviewquestion.com/blog/advanced-rag-patterns-interview-guide).

How to choose between them

Follow-ups missing context: contextualize. Vocabulary mismatch: multi-query plus RRF, or HyDE for zero-shot dense retrieval. A very specific question inside a broader topic: step-back. Independent parts: parallel decomposition. Hop two needs hop one: sequential decomposition.

The short version: step-back moves up, decomposition splits sideways or down, multi-query rephrases, and HyDE changes the shape of the query.

Latency and cost

Every rewrite adds at least one LLM call before retrieval. In query2doc's measurement, BM25 index search took 16 ms, while with query2doc the LLM call took over 2000 ms and search took 177 ms, because longer expanded queries also slow the inverted index. That is a 2023 figure for an API whose latency "depends on server load", not a standard. RAG-Fusion measured 1.77 times the runtime.

Cheaper options from the sources: combine the rewrite with the retrieval check in one call, use a smaller or fine-tuned rewriter, run variant retrievals in parallel (LangChain's async path, LlamaIndex's use_async), and skip retrieval when the check returns false. For a per-stage budget, see RAG latency and cost (https://aiinterviewquestion.com/blog/rag-latency-and-cost-interview-guide).

Failure modes interviewers probe

The follow-up is searched as-is, with no history. Fix: contextualize first.

A variant drifts from intent and fusion rewards the wrong documents. Fix: keep the original query, generate fewer variants, rerank against the original question, and log every generated query.

HyDE's pseudo-document carries false specifics that retrieval follows. Fix: keep the original query in the embedding.

Results are merged by plain union, so order is arbitrary and, with LangChain's default, the original query is not searched. Fix: RRF or a reranker, plus include_original=True.

A multi-hop question gets one-shot retrieval, so hop-two facts are never fetched. Or a bad sub-answer flows into synthesis unchecked.

A good rewrite still retrieves the wrong thing, as in Step-Back's 45% retrieval errors, so evaluate retrieval separately.

An illustration, not a sourced incident: for "How long does the return policy cover?", one variant becomes "warranty coverage duration", RRF lifts warranty pages into the top-k, and the answer quotes the warranty period instead of the return window.

How to evaluate a query rewriter

Score retrieval for rewritten versus original queries on a labeled set (recall@k or hit rate), then score answers end to end. Keep both: Step-Back's analysis shows abstraction, retrieval, and reasoning errors need different fixes. See How to debug a RAG pipeline (https://aiinterviewquestion.com/blog/how-to-debug-a-rag-pipeline-retrieval-ranking-and-grounding).

Log every generated query and review samples by hand; RAG-Fusion's off-topic answers traced back to weak generated queries. Track the latency each stage adds. No source gives a standard rewrite metric or variant count; set both from your eval set. Practice: Evaluate query rewriting quality in an AI search product (https://aiinterviewquestion.com/questions/evaluate-query-rewriting-quality-in-an-ai-search-product).

Follow-up questions to prepare

Why can HyDE hurt? (False details, the same-semantics assumption, and small drops with a weaker LLM over a fine-tuned encoder.) Why not just union the results? (Union does not rank.) When is step-back better than decomposition? (When one broader retrieval covers the specific question.) How do you keep rewriting cheap? (One combined call, a smaller model, parallel retrieval, and no search when none is needed.) What if a multi-query rollout makes results worse? (Read the logged variants, compare retrieval with and without the original query, and check the merge. See Debug a multi-query retrieval regression (https://aiinterviewquestion.com/questions/debug-multi-query-retrieval-regression-in-a-multi-tenant-ai-support-platform).)

Sample interview answer

"I start by finding why retrieval misses. If follow-ups fail, I add a contextualization step that rewrites the turn into a standalone query, combined with the check for whether the turn needs retrieval, so it costs one small-model call, as in OpenAI's latency guide. For vocabulary mismatch I try multi-query retrieval, keeping the original query in the set because LangChain's default leaves it out. I merge with Reciprocal Rank Fusion, k = 60, then rerank against the original question, because a plain union does not rank. For short questions over a dense index with no labels, HyDE is an option: embed a hypothetical answer plus the query, and never show that text to users, since it can contain factual errors. For multi-hop questions I decompose, in parallel when the parts are independent and sequentially when hop two needs hop one, because IRCoT showed one-step retrieval is not enough for multi-step QA. I don't assume any of this helps. RAG-Fusion's case study saw off-topic answers when variants drifted, and Step-Back's analysis found 45% of errors were retrieval errors even with an on-target step-back question. So I evaluate the rewriter on its own: recall@k for rewritten versus original queries, end-to-end quality, added latency, and logged queries reviewed for drift. I tune the variant count on that set."

The close: rewrite the search input, not the user's question; merge and rerank against the original; and measure recall, drift, and latency before you ship a rewriter.

https://arxiv.org/abs/2305.14283 https://developers.openai.com/api/docs/guides/latency-optimization https://reference.langchain.com/python/langchain-classic/retrievers/multi_query/MultiQueryRetriever https://plg.uwaterloo.ca/~gvcormac/cormacksigir09-rrf.pdf https://arxiv.org/abs/2402.03367 https://arxiv.org/abs/2212.10496 https://arxiv.org/abs/2303.07678 https://arxiv.org/abs/2310.06117 https://arxiv.org/abs/2205.10625 https://arxiv.org/abs/2210.03350 https://arxiv.org/abs/2212.10509 https://developers.llamaindex.ai/python/framework/optimizing/advanced_retrieval/query_transformations/ https://developers.llamaindex.ai/python/examples/query_engine/sub_question_query_engine/

RAGQuery rewritingHyDEQuery decompositionMulti-query retrievalQuery expansionStep-back prompting

Questions in this guide

Deep explanations with architecture diagrams for every question below.