Skip to main content
AI Interview Question
INTERVIEW GUIDERAG8 questions7 min readOct 4, 2026

When to add a reranker: recall, precision, and the latency/cost interview answer

Retrievers own recall; cross-encoders reorder a shortlist. Cite BEIR carefully, label vendor cost claims, and diagnose missing recall before you add latency.

When to add a reranker: recall, precision, and the latency/cost interview answer

Interviewers asking when to add a reranker want a clean two-stage story. First-stage retrieval—lexical search such as BM25, a bi-encoder dense index, or a hybrid of both—owns recall. Its job is to put relevant evidence into a candidate set of size K. Second-stage reranking—usually a cross-encoder or a managed rerank API—owns precision and ordering inside that set. It promotes useful candidates into the small N passages you send to the LLM. Sentence Transformers documents the canonical retrieve-and-re-rank pipeline: retrieve a large shortlist (their example is about 100 hits) with lexical search or a bi-encoder, then score those pairs with a CrossEncoder that attends over query and document together. Cohere frames the same pattern for production RAG: keep keyword or vector search first, call Rerank on that list, then take top_n into generation. NVIDIA describes the same workflow—embeddings narrow millions down to tens of candidates, then a reranker produces a final small set, often about five passages. The interview one-liner is simple: the retriever owns recall; the reranker owns ordering inside the candidate set. If recall is broken, fix retrieval. Do not bolt on a cross-encoder.

Say the bi-encoder versus cross-encoder distinction the way Sentence Transformers does. A bi-encoder encodes the query and the document independently, compares embeddings with cosine or dot similarity, and scales because documents can be precomputed and indexed. A cross-encoder concatenates query and document, runs joint attention, and outputs a single relevance score. Pairwise performance is generally stronger, but each pair needs a forward pass, so a cross-encoder is too slow to score thousands or millions of documents and is used to re-rank top-k from a first-stage retriever. Cross-encoders do not produce indexable sentence embeddings for corpus-wide approximate nearest-neighbor search. If the interviewer goes deeper, late-interaction models such as ColBERT sit between dense bi-encoders and full cross-encoders—token-level interaction with precomputable document representations. BEIR evaluated ColBERT as a late-interaction system with strong results and higher compute and index cost than pure dense retrieval. Do not equate ColBERT with a classic cross-encoder reranker.

Cite BEIR carefully. Thakur et al. built a heterogeneous zero-shot IR benchmark across 18 datasets and 9 task types, with nDCG@10 as the primary metric in their experiments. Their abstract states that re-ranking and late-interaction models on average achieve the best zero-shot performances, at high computational cost. The architecture they call BM25+CE runs BM25 first, then a MiniLM cross-encoder re-ranks the top-100 hits. Among the ten systems they tested, BM25+CE was their best overall zero-shot system: it beat BM25 on 16 of 18 datasets, and the average-performance-versus-BM25 row in their Table 2 summary is +11% relative. Those gains are zero-shot IR averages across BEIR, not a claim that every production RAG product needs a reranker. On ArguAna, BM25 nDCG@10 is 0.315 and BM25+CE is 0.311—the cross-encoder did not help. On Touché-2020, BM25 is 0.367 and BM25+CE is 0.271—the cross-encoder hurt. Illustrative wins elsewhere in that table include TREC-COVID 0.656 to 0.757, HotpotQA 0.603 to 0.707, FEVER 0.753 to 0.819, and SciFact 0.665 to 0.688. BEIR also stresses that strong in-domain MS MARCO scores do not predict zero-shot generalization, and BM25 remains a robust baseline. Their DBPedia 1M latency probe puts BM25+CE at about 450 ms on GPU and about 6100 ms on CPU estimated retrieval latency, against BM25 at about 20 ms on CPU—best quality in their set, much slower than dense-only retrieval under about 20 ms on GPU in their table.

Latency and cost are two different conversations. Cross-encoder work scales with candidate count K and passage length because the model scores joint sequences. Prefer batching all K pairs, capping K from recall curves on labeled queries, measuring p50, p95, and p99, and defining a timeout that falls back to first-stage order. Do not invent universal latency bands; measure your stack. Cost has two sides. Reranker compute or API units grow with K times tokens. LLM generation cost often dominates end-to-end spend. NVIDIA argues a reranker can reduce end-to-end RAG cost by sending fewer chunks to the LLM while keeping or improving accuracy. For their NeMo Retriever Llama 3.2 reranking NIM benchmarks, they publish 21.54% cost savings driven mainly by fewer LLM input tokens, and they state that in their pricing comparison Llama 3.1 8B costs roughly 75 times more than their Llama 3.2 reranker to process five chunks and generate an answer. Those figures are vendor-specific NeMo and NIM comparison claims from their March 2025 technical blog—not industry law. Their useful variables are N_Base chunks to the LLM without rerank, N_Reranked after rerank, and K candidates scored by the reranker. Require N_Reranked less than or equal to N_Base for the same-or-better context-budget story, and trade maximizing accuracy against maximizing savings.

Hybrid retrieval and reranking compose; they are not substitutes. Hybrid—dense plus BM25 or reciprocal rank fusion—improves candidate recall. Reranking improves ordering of that pool. The production chain is hybrid or dense or lexical retrieval, then top-K into the reranker, then top-N into the LLM. Cohere notes that Rerank works on semantic or lexical first-stage results. If candidate recall is weak on both keywords and paraphrases, strengthen hybrid or retrieval before spending more on reranking. Keep two K’s distinct: candidate K into the reranker from Recall@K curves, and final N into context from evidence coverage and usable context—not from the model’s marketed max context window. Managed APIs add provider limits, not physics. Cohere recommends against sending more than 1,000 documents in a single Rerank request, and the API hard-errors if documents times max_chunks_per_doc exceeds 10,000. Long documents may be truncated or auto-chunked. Relevance scores in [0, 1] are for ranking; 0.9 is not twice as relevant as 0.45.

Add or test a second-stage reranker when candidate recall is already acceptable—relevant evidence often appears in top-K—but precision@N, nDCG@N, or MRR for the context window is weak and the right docs sit below noise. Add it when queries are semantically hard: paraphrase, multi-aspect, near-duplicates, or semi-structured fields where joint query–passage attention helps. Cohere documents multi-aspect, YAML-structured, tabular, and multilingual use cases for their Rerank models. Add it when better ranking lets you reduce N sent to the LLM and still hit answer quality—that is the lever NVIDIA’s cost story depends on. Add it when you can afford the p95 latency and have a timeout fallback to original order or a cheaper model. And add it only if you will ablate on your labeled queries: retriever-only versus retriever plus rerank versus improved retrieval without a reranker.

Skip or defer a reranker when recall is the bottleneck and relevant chunks rarely enter top-K because of chunking, filters, embedding mismatch, or a missing hybrid. A reranker cannot recover them. Skip when first-stage ranking already saturates your gate with high precision@N on the evaluation set. Skip when the workload is mostly exact ID or keyword lookups that lexical search already ranks well. Skip when the latency SLO cannot absorb second-stage cost—BEIR’s latency gap is qualitative evidence; measure yours. Skip when the failure is generation: the model ignores or misquotes retrieved context—fix prompting, citations, or grounding checks, not ranking. Skip when you cannot label or observe stage-level metrics, or when permissions and PII cannot be applied before an external rerank API. Skip when the corpus is tiny enough that a simpler strategy—even a cross-encoder over the whole set for in-document search, as Sentence Transformers notes—or plain retrieval is enough.

The critical failure mode you must teach is that a reranker cannot recover missing first-stage recall. With a fixed candidate set of size K, Recall@K is unchanged by reordering. Only metrics sensitive to order inside the set—nDCG@N, MRR, Precision@N, Recall at final N—can improve. Sentence Transformers’ NanoBEIR CrossEncoder evaluator encodes the same idea: if positives are not already in the BM25 top rerank_k, realistic evaluation is bound by first-stage recall when always_rerank_positives is false. Close the interview with a diagnosis from logs. Was the gold chunk in top-K before rerank? If no, fix retrieval. If yes but not in top-N after, fix ranking or N. If it is in top-N but the answer is wrong, fix generation.

What interviewers expect fits a short checklist. Define bi-encoder versus cross-encoder and why the cross-encoder only sees top-K. Name candidate K versus context N and how you choose them from recall curves and latency. State the recall ceiling: rerank reorders; it does not expand the corpus hit set. Mention hybrid as a recall lever that pairs with rerank as a precision lever. Discuss latency and cost honestly: the cross-encoder adds compute and may reduce LLM tokens if N shrinks; cite vendor numbers only as vendor claims. Propose an ablation with Recall@K, nDCG@N or MRR, answer faithfulness, p95 latency, and cost per thousand queries. Know BEIR as evidence that BM25 plus a cross-encoder often lifts zero-shot nDCG@10 on average, with exceptions and cost. Do not claim always plus X percent accuracy without measuring on the company’s corpus.

https://www.sbert.net/examples/sentence_transformer/applications/retrieve_rerank/README.html https://sbert.net/docs/cross_encoder/usage/usage.html https://arxiv.org/abs/2104.08663 https://arxiv.org/html/2104.08663v4 https://github.com/beir-cellar/beir https://developer.nvidia.com/blog/how-using-a-reranking-microservice-can-improve-accuracy-and-costs-of-information-retrieval/ https://docs.cohere.com/docs/reranking-with-cohere https://docs.cohere.com/docs/reranking-best-practices https://docs.cohere.com/reference/rerank https://ctxwire.com/articles/rag-reranking-guide/ https://docs.pinecone.io/guides/search/rerank-results

RerankingCross-encoderBi-encoderRAGLatencyPrecision

Questions in this guide

Deep explanations with architecture diagrams for every question below.