Skip to main content
AI Interview Question
INTERVIEW GUIDERAG8 questions4 min readOct 4, 2026

Picking an embedding model for retrieval

Choose the model by retrieval task and corpus. Pair the query encoder with the document encoder. Swap the function and you rebuild the index.

Picking an embedding model for retrieval

Interviewers asking how you pick an embedding model for retrieval want the decision framed by task and corpus, not by a leaderboard average. Dense Passage Retrieval uses two independent BERT networks (base, uncased), one question encoder and one passage encoder. Each outputs a 768-dimensional vector from the [CLS] token, and the score is the dot product of those two vectors. In the OpenAI guide’s code-search example, functions are indexed with text-embedding-3-small and the query is embedded with that same model. That example is not a rule that every vendor uses one shared encoder. Voyage documents that voyage-4-large, voyage-4, and voyage-4-lite embeddings are compatible with each other. That compatibility is a Voyage family claim. Do not generalize it across vendors. Swapping model, dimension, pooling, normalization, or instruction template means the index must be rebuilt.

MTEB is eight task types, not only retrieval. On the English subset Table 1, ST5-XXL has the highest average at 59.51, with retrieval 42.24 and STS 82.63. SGPT-5.8B-msmarco averages 58.81, with retrieval 50.25 and STS 78.10. SimCSE-BERT-sup scores 79.12 on STS and 21.82 on retrieval. The paper’s own lines are that no method dominates all tasks, that STS-geared models perform badly on retrieval, and that retrieval is asymmetric. The main retrieval metric is nDCG@10.

MMTEB, Enevoldsen et al., arXiv 2502.13595, ICLR 2025, reports multilingual-e5-large-instruct at Borda rank 1 on MTEB(Multilingual), average across all tasks 63.2 and retrieval 57.1, with 560 million parameters named in the abstract. GritLM-7B is rank 2, average across all tasks 60.9 and retrieval 58.3. Those are the paper’s results. They are not an October 2026 live leaderboard fact. No model named here is claimed as a current number one.

In-domain score is a weak signal for out-of-distribution transfer. Thakur et al. write that in-domain performance is not a good indicator for out-of-domain generalization, and that no single approach consistently outperforms other approaches on all datasets. Use that transfer point in the interview. Dense-versus-BM25 comparisons and reranker design belong in separate posts.

Dimension choice is not free chopping. Kusupati et al. optimize the first m dimensions as nested representations. Do not bless truncating an arbitrary embedding to a shorter width and calling it Matryoshka. Post-hoc methods such as SVD drastically lose accuracy relative to MRL as size decreases. The paper’s 14× smaller and faster figures are ImageNet results, not text RAG. On the OpenAI side, text-embedding-3-small defaults to 1536 dimensions and text-embedding-3-large to 3072, both with a maximum input of 8192 tokens, and both allow a dimensions parameter. Large shortened to 256 still beat unshortened ada-002 at 1536 on MTEB. If you slice after the fact, L2-normalize. Vendor MTEB figures are 62.3% for small, 64.6% for large, and 61.0% for ada-002. Do not invent prices.

Asymmetric and instruction-prefixed models add a template you have to keep. E5-base-v2 needs the query: and passage: prefixes or performance degrades. It is 768-dimensional, English, and truncates at 512 tokens. The card’s example normalizes embeddings with L2 and then scores with a matrix product. Qwen3-Embedding-0.6B puts the instruction on the query only. The vendor claim is that dropping the instruct drops retrieval by about 1% to 5%. It supports MRL and is apache-2.0. Voyage prepends its own prompts when input_type is query or document.

Context length is the model’s token limit, not the chunker’s character count. E5-base-v2 is 512. OpenAI text-embedding-3 models are 8192; the guide does not say whether overflow errors or truncates. The Voyage-4 family is 32,000 tokens, voyage-law-2 is 16,000, and truncation defaults to True. The Qwen card conflicts between 32k and a Transformers example max_length of 8192. Do not pick one of those numbers in the interview without checking the card you will ship.

The similarity metric follows the model card. OpenAI recommends cosine similarity. Its embeddings are normalized to length 1, so cosine can be computed with a dot product. The FAQ’s identical-rankings line is between cosine and Euclidean distance, not between cosine and dot product. DPR trains for inner product. The E5 and Qwen examples normalize with L2 and then score with a matrix product. Cosine on unnormalized vectors is a different ranking than those cards specify.

API versus local hosting is a privacy and license choice. Do not invent prices, and do not state a license the card does not name.

The interview close is simple. Pick by the retrieval task and the corpus you will actually index. Accept that changing the embedding function means rebuilding that index.

https://arxiv.org/html/2004.04906 https://developers.openai.com/api/docs/guides/embeddings https://docs.voyageai.com/docs/embeddings https://arxiv.org/html/2210.07316v2 https://arxiv.org/abs/2502.13595 https://arxiv.org/html/2104.08663v4 https://arxiv.org/html/2205.13147v4 https://huggingface.co/intfloat/e5-base-v2 https://huggingface.co/Qwen/Qwen3-Embedding-0.6B

EmbeddingsDense retrievalMTEBMatryoshka

Questions in this guide

Deep explanations with architecture diagrams for every question below.