Design Hybrid Search architecture for a coding copilot for a large engineering org
Mid-Level architecture interview question on Hybrid Search within RAG.
Read full explanationSeparate BM25 from dense retrieval, then fuse with reciprocal rank fusion. Cite the datasets where each wins. Do not average raw scores with cosine.

Interviewers asking about dense versus hybrid search want three separations. Lexical search, usually BM25, is weighted term matching. Dense search uses two encoders and similarity in a shared vector space. They fail on different queries. Do not average raw BM25 scores with cosine or dot-product scores. Hybrid means retrieve from both, then fuse. The usual fuse is Reciprocal Rank Fusion, not a weighted average of raw scores.
BM25 is the sparse lexical retriever. Robertson and Zaragoza, The Probabilistic Relevance Framework: BM25 and Beyond (2009), is the standard reference. Sciavolino et al., EMNLP 2021, describe it as weighted term matching that is not trained on one query distribution. It is strong at lexical match and fails on synonyms and paraphrases.
Dense passage retrieval is Karpukhin et al., EMNLP 2020. Two BERT encoders are trained contrastively, passages are pre-indexed, and the system returns the top passages by similarity. Karpukhin’s Figure 1 says a DPR model trained on 1,000 examples already outperforms BM25, measured on the Natural Questions development set. Sciavolino cites that result.
The fusion to name is Reciprocal Rank Fusion, from Cormack, Clarke, and Buettcher, SIGIR 2009. The score is the sum, over the rankings that contain the document, of 1/(k + rank). A pilot found k=60 near-optimal, and not critical. On Elastic’s top-N lists, a document missing from one list contributes nothing for that list.
In-domain natural questions, dense wins. Sciavolino et al., arXiv v3 Table 1, reports top-20 accuracy of 80.1 for DPR trained on Natural Questions, against 64.4 for BM25. The EMNLP proceedings table prints that BM25 cell as 64.5. Those figures are top-20 accuracy on one Wikipedia setup from 2021. They are not a 2026 production RAG number.
The same comparison flips on EntityQuestions, simple Wikidata facts such as “Where was Arve Furset born?” On the arXiv v3 table, top-20 accuracy is 49.7 for DPR, 56.7 for a multi-dataset DPR, and 72.0 for BM25. The proceedings table prints that BM25 average as 71.2. On “Where was [E] born?” the arXiv gap is 25.4 versus 75.3, and the proceedings table prints 75.2. DPR holds on common entities and collapses on rarer ones unless that pattern was in training.
Zero-shot, BM25 is the baseline to beat. Thakur et al., BEIR (2021), cover 18 datasets. BM25 underperforms neural methods by 7–18 nDCG@10 points in-domain on MS MARCO, then is a strong zero-shot baseline. Their DPR zero-shot average is 47.7% below BM25. No single approach consistently outperforms the others on every dataset. BM25 plus a cross-encoder was +11% versus BM25 on the average row and outperformed BM25 on 16 of 18 datasets. It fails on ArguAna and Touché-2020.
Hybrid is a small, stable gain when the two retrievers are already close, not a guaranteed jump. Elastic, on 20 July 2023, fused BM25 with the Elastic Learned Sparse Encoder using reciprocal rank fusion, on BEIR. ELSER is learned sparse. It is not BM25, and it is not a dense bi-encoder. SPLADE is the same family, and these sources do not include a SPLADE score. That fusion raised average nDCG@10 by 1.4% over the sparse encoder alone and by 18% over BM25 alone, and it was better or similar to BM25 on every set they tested. Those figures are ELSER plus BM25 on Elastic’s BEIR slice, not a claim that hybrid always adds 18%. For the dense models they grid-searched, the difference between the best and worst k and N combinations was only about 5%. A calibrated linear mix of normalized scores did better in the best case, +6% versus the sparse encoder and +24% versus BM25, but the best weight did not transfer. On ArguAna it was possible to beat RRF with about 40 annotated queries, and that threshold varied by dataset. RRF is the plug-and-play choice. Weighted scores are not.
Index size can flip the winner. Reimers and Gurevych (arXiv:2012.14210) found that on MS MARCO a dense model trained without hard negatives beat BM25 from 10k through 1M passages and lost at the full 8.8M. Table 1 reports dev MRR@10 times 100: 17.34 for the 768-dim model without hard negatives and 17.56 for BM25 at 8.8M passages. The same architecture trained with hard negatives still led at 8.8M, 28.55 versus 17.56. Do not say dense dies at scale. The gap closes, and a weakly trained dense model can fall behind BM25 once the index is large.
Hybrid costs two indexes and two queries. Elastic ran them sequentially, and BM25 was the faster of the two. BEIR Table 3, on 1M DBPedia documents, puts BM25 at about 20ms on CPU with a 0.4GB index. A 768-dim dense model was about 14–20ms on GPU, 125–275ms on CPU, and a 3GB index. BM25 scores are not on the same scale as cosine or as ELSER scores. Normalize before any weighted sum. Reciprocal rank fusion uses ranks and sidesteps that.
A reranker cannot recover a passage the first stage never returned. Hybrid is a recall move. BEIR’s best average system was still BM25 plus a cross-encoder, a second stage, not a substitute for the first. Skip hybrid only when you have measured that the query mix is one-sided, or the latency budget cannot take the second retriever. Do not skip it because a demo felt semantic.
https://aclanthology.org/2021.emnlp-main.496/ https://aclanthology.org/2020.emnlp-main.550/ https://arxiv.org/html/2109.08535v3 https://cormack.uwaterloo.ca/cormacksigir09-rrf.pdf https://arxiv.org/abs/2104.08663 https://www.elastic.co/search-labs/blog/improving-information-retrieval-elastic-stack-hybrid https://arxiv.org/abs/2012.14210
Deep explanations with architecture diagrams for every question below.
Mid-Level architecture interview question on Hybrid Search within RAG.
Read full explanationMid-Level evaluation interview question on Dense Retrieval within Vector Databases.
Read full explanationMid-Level trade-off interview question on BM25 within Vector Databases.
Read full explanationMid-Level debugging interview question on Hybrid Search within Vector Databases.
Read full explanationMid-Level architecture interview question on BM25 within Vector Databases.
Read full explanationMid-Level scenario interview question on Hybrid Search within Vector Databases.
Read full explanationMid-Level implementation interview question on Dense Retrieval within Vector Databases.
Read full explanationStaff evaluation interview question on Access Control within RAG.
Read full explanation