Skip to main content
AI Interview Question
INTERVIEW GUIDERAG8 questions5 min readOct 3, 2026

Chunk size and overlap in RAG: what interviewers expect

Interviewers want chunk size and overlap trade-offs, not a magic default. Cover recursive vs fixed, tokens vs characters, and how you evaluate.

Chunk size and overlap in RAG: what interviewers expect

When an interviewer asks how you choose chunk size and overlap for production RAG, they are not looking for a magic number. They want you to reason about precision, context, cost, and how you would evaluate the choice on your own corpus. Vendor defaults and blog examples are starting points. They are not a universal answer.

Chunk size is a three-way trade-off. Chunks that are too small can raise precision for factoid lookup, but they lose surrounding context, multiply the number of embeddings you store, and often return incomplete or near-duplicate hits. Chunks that are too large dilute the topic signal in the embedding, can bump into the embedding model's input limit, waste LLM context with irrelevant text, and make precise citation harder. There is no single size that wins for every document shape, query type, embedding model, top-k setting, or reranker setup. Treat published defaults as experiments you still have to run.

Evidence of that corpus dependence shows up in NVIDIA's June 2025 study across multiple datasets. Extreme sizes such as 128 and 2,048 tokens often underperformed medium sizes. Factoid-heavy sets favored smaller or medium chunks around 256 to 512 tokens, while more analytical queries benefited from larger chunks around 1,024 tokens or from page-level boundaries. Page-level chunking was a strong default in their PDF-heavy setup. That does not mean page-level chunking is best for every corpus. The study is at https://developer.nvidia.com/blog/finding-the-best-chunking-strategy-for-accurate-ai-responses/

Overlap exists to reduce boundary loss. When a fact, sentence, or clause straddles two chunks, overlap raises the chance that at least one retrieved chunk still contains the full unit. The cost is real: more chunks mean higher embed and storage cost, and adjacent near-duplicates can crowd top-k and hurt diversity. Interviewers like hearing that trade-off.

Vendor numbers are defaults to evaluate, not proven optima. Azure AI Search guidance often starts fixed-size chunking at about 512 tokens with roughly 25 percent overlap (128 tokens), and also discusses about 10 to 15 percent overlap as a common fixed-size pattern. See https://learn.microsoft.com/en-us/azure/search/vector-search-how-to-chunk-documents and related notes in https://learn.microsoft.com/en-us/azure/search/cognitive-search-skill-textsplit and https://learn.microsoft.com/en-us/azure/architecture/ai-ml/guide/rag/rag-chunking-phase. NVIDIA tested 10 percent, 15 percent, and 20 percent overlap and reported 15 percent best on FinanceBench with 1,024-token chunks. That was not a full grid search across every dataset. Overlap is most useful for fixed or recursive windows on prose. If you split on strong structure such as headers, records, or AST code units, overlap is often less necessary. Still evaluate before you drop it.

Fixed-size splitting cuts every N characters or tokens, optionally with overlap. It is simple and predictable, but without boundary logic it can cut mid-sentence or mid-code. Recursive splitting, as in LangChain's RecursiveCharacterTextSplitter, tries separators in order. The default order is paragraphs, then lines, then spaces, then characters, so larger units stay together as long as possible. chunk_size is a max under the chosen length_function, and chunk_overlap mitigates boundary loss. Official LangChain docs measure size by characters unless you pass a token length function. Docs: https://docs.langchain.com/oss/python/integrations/splitters/recursive_text_splitter. A strong interview answer prefers structure-aware recursive or Markdown-aware splitting over naive fixed cuts for prose, while still enforcing a max size the embedding model can accept.

Semantic chunking places breakpoints using embedding similarity between sentences so each unit is more topically coherent. LlamaIndex documents this pattern with SemanticSplitterNodeParser and related node parsers. Caveats interviewers expect you to name: English-centric sentence regexes, breakpoint percentile tuning, extra embedding cost at ingest, variable chunk lengths, and the fact that semantic splitting is not automatically better than recursive splitting plus evaluation on your data. Primary write-ups: https://developers.llamaindex.ai/python/examples/node_parsers/semantic_chunking/ and https://developers.llamaindex.ai/python/framework/module_guides/loading/node_parsers/modules/

Embedding and chat models count tokens, not characters. A character-based splitter can produce chunks that exceed the embedding model's token limit or undershoot the size you thought you set. OpenAI's cookbook on long inputs stresses that embedding models have a token max context, that long inputs must be truncated or chunked by tokens, and that splitting on paragraph or sentence boundaries is preferred when possible. See https://developers.openai.com/cookbook/examples/embedding_long_inputs. Production practice is to size and overlap with the same tokenizer family as the embedding model, or with the vendor's token unit such as Azure's azureOpenAITokens. Do not conflate character length with token length in an interview answer.

Boundary awareness matters for code and Markdown. Prose separators are wrong for source files. Prefer language-aware separators, such as RecursiveCharacterTextSplitter.from_language, or AST and syntax-aware splitters so you do not cut mid-function or mid-block. LangChain's code splitter notes are at https://docs.langchain.com/oss/python/integrations/splitters/code_splitter For Markdown or HTML, split on headings or sections first, keep the heading path in metadata, then recursively split oversized sections. Overlap typically applies within a section, not across header boundaries. That cross-boundary overlap is a common pitfall when you chain a Markdown header split into a recursive split. Tables and charts should stay intact when possible. NVIDIA notes that tables and charts were not token-split in their pipeline.

To evaluate chunk configs, change one variable at a time: size, then overlap, then splitter type. Use retrieval metrics such as Recall at k, Precision at k, MRR, and nDCG on a labeled query set, plus end-to-end answer quality. NVIDIA used RAGAS-style answer accuracy with LLM judges in their study. Qualitative checks still matter: inspect boundary failures where an answer span is cut in half, duplicate near-neighbors in top-k, and whether citations remain usable. Re-evaluate when the corpus, embedding model, or query mix changes. Chunking is pipeline-specific.

A strong interview answer sounds like this. First state the trade-off among precision, context, and cost. Then name a sensible starting default and why, for example recursive splitting with token counting, roughly 256 to 512 tokens for factoid work and larger for analytical queries, labeled clearly as a start. Explain overlap as boundary insurance with cost and diversity downsides. Prefer structure such as headers, code units, or pages when documents have it. Insist on offline evaluation on your own queries before locking production settings. Cite the sources you actually used rather than inventing a universal best number.

ChunkingOverlapRAGRetrievalEmbeddings

Questions in this guide

Deep explanations with architecture diagrams for every question below.