Debug Faithfulness regression in a multi-tenant AI support platform
Senior debugging interview question on Faithfulness within RAG.
Read full explanationRAG model ignoring retrieved context? Prove the gold chunk reached the prompt, find whether memory or a wrong chunk won, then fix the layer that failed.

"The model ignored the retrieved context" is one of the most common RAG complaints, and interviewers use it to see whether you debug or guess. The weak answer is "add 'only use the context' to the prompt". The strong answer first proves the answer was actually in the prompt, then names which of several different failures happened, then fixes the layer that caused it: prompt, context assembly, decoding, training, or verification.
Put simply: retrieval brought the right page, and the model answered from memory anyway, or from the wrong part of the page, or answered when the page did not support any answer. Each of those has a different fix. This guide picks up where the grounding section of How to debug a RAG pipeline (https://aiinterviewquestion.com/blog/how-to-debug-a-rag-pipeline-retrieval-ranking-and-grounding) stops.
First prove the context was there
Before you blame the model, open the exact prompt that was sent for the failing request. Is the answer-bearing chunk in it, complete, and readable? If the chunk never made it in, this is a retrieval or ranking miss, and the work is in When to add a reranker (https://aiinterviewquestion.com/blog/when-to-add-reranker-rag-interview) or Chunking and overlap (https://aiinterviewquestion.com/blog/chunking-and-overlap-rag-interview-guide), not in the generator. If a split chunk holds only half the fact, the model did not ignore anything. It had half the evidence.
Only when the gold chunk is in the final prompt, and the answer still contradicts it or skips it, is this a generation failure. Say that boundary out loud in an interview. It is the step most candidates skip.
Two opposite failures look the same
The ticket says "it ignored the context", but there are two failures that point in opposite directions.
Prior wins over correct context. The model has a memorized answer and keeps it even though the passage says something else. Longpre et al. (EMNLP 2021) called this a knowledge conflict and tested it by swapping the answer entity inside the passage for another entity. Their memorization ratio counts how often the model still outputs the original, memorized answer instead of the one the passage now supports. They found over-reliance on memorized information, and the ratio rose with model size across T5 models from 60M to 11B parameters. Their conclusion for practitioners: evaluate whether the model tends to hallucinate rather than read.
Context wins over a correct prior. The opposite case: the retrieved text is wrong, outdated, or perturbed, and the model repeats it. In ClashEval (2024), Wu, Wu, and Zou perturbed answers in content for 1,294 questions across six domains such as drug dosages and Olympic records, and tested six models including GPT-4o, Claude Opus, and Gemini 1.5. The models adopted incorrect retrieved content, overriding their own correct prior knowledge, over 60% of the time. The more unrealistic the edit, the less likely the model was to adopt it. The less confident the model was in its own first answer, measured by token probabilities, the more likely it was to take the context.
Xie et al. (ICLR 2024) found both behaviors in one study. LLMs can be highly receptive to conflicting evidence when it is coherent and convincing, yet show strong confirmation bias when the evidence also contains something that agrees with their memory. That bias was stronger for more popular facts.
The interview point: "follow the context" is not always the right goal. You want the model to follow correct context and resist wrong context, so you need to know which side failed.
Three more ways context gets skipped
Position. Liu et al., Lost in the Middle (TACL), found performance is often highest when the relevant information is at the beginning or end of the input and drops when it sits in the middle, even for long-context models. A gold chunk ranked 7th of 12 can be present and still under-used.
Distractors. Freda Shi et al. (ICML 2023) built GSM-IC, grade-school math problems with irrelevant sentences added, and found model accuracy dropped dramatically. Yoran et al. (2023) analyzed five open-domain QA benchmarks and characterized the cases where adding retrieval reduces accuracy, with misused irrelevant evidence causing cascading errors in multi-hop questions. Padding the prompt with near-misses is not neutral.
Insufficient context, answered anyway. Joren et al. (ICLR 2025, a UC San Diego, Duke, and Google team) asked whether RAG errors come from the model failing to use the context or from context that cannot answer the question. They built a sufficient-context autorater with Gemini 1.5 Pro at 93% accuracy. Larger models such as Gemini 1.5 Pro, GPT-4o, and Claude 3.5 did well when the context was sufficient but often gave a wrong answer instead of abstaining when it was not. Smaller models such as Mistral 3 and Gemma 2 hallucinated or abstained often, even with sufficient context. Here the bug is not ignoring the context. It is answering past it.
How to diagnose it
Build a small frozen set from real failures, then run three cheap tests per item. These are practice built on the papers' methods, not benchmark results.
Closed-book versus open-book. Ask the same question with no context and with the retrieved context. If both answers match and both are wrong, the model is answering from its prior. If the closed-book answer was right and the open-book answer is wrong, a bad chunk overrode it. ClashEval names these two rates prior bias (the model uses its prior when the context was right) and context bias (the model uses the context when the context was wrong and the prior was right). Track both. Pushing one down can push the other up.
Counterfactual swap. Copy a gold chunk, change the key fact the way Longpre et al. swapped answer entities, and ask again. A context-faithful system follows the edited passage. A stubborn one returns the memorized fact. This tells you whether the model reads, not whether the corpus is right.
Position and noise swap. Move the gold chunk to the first slot, then add or remove distractor chunks. If accuracy changes, it is an ordering or context-size problem, and the fix is in assembly, not instructions.
Log a sufficiency label on each item: does the context actually answer the question? An answer to an insufficient context should be an abstention, and should be scored that way.
Fixes by layer
Prompt. Anthropic's hallucination guide recommends explicitly allowing "I don't know", restricting answers to the provided documents, and for long documents (over 20k tokens) extracting word-for-word quotes first and answering from them. Zhou et al. (EMNLP 2023 Findings) found two prompting strategies most effective for context faithfulness without training: opinion-based prompts, which reframe the passage as a narrator's statement and ask for the narrator's opinion (Bob said, "..." Q: ... in Bob's opinion?), and counterfactual demonstrations, which use few-shot examples containing false facts to improve faithfulness when the passage and memory conflict. Freda Shi et al. found an instruction to ignore irrelevant information also helped on GSM-IC.
Context assembly. Prefer fewer, better chunks: the distractor and position results above both argue against padding. Rerank, put the strongest evidence where the model uses it, and cut near-duplicates. For long multi-document inputs (20k+ tokens), Anthropic recommends putting the documents at the top, above the query and instructions, and reports that queries at the end can improve response quality by up to 30% in its tests. It also recommends wrapping each document in tags with its source. Yoran et al. showed an NLI filter that drops passages which do not entail the question and answer prevents the accuracy loss, but also throws away relevant passages, so measure recall when you add one. The cost side of trimming context is in RAG latency and cost (https://aiinterviewquestion.com/blog/rag-latency-and-cost-interview-guide).
Decoding. Context-aware decoding (CAD) by Weijia Shi et al. (2023) samples from a contrast of two distributions: the model's logits with the context, weighted up, minus its logits without the context. That amplifies what the context changes. Without extra training it gave LLaMA-30B a 14.3% gain on a summarization factuality metric and a 2.9x improvement on Longpre's knowledge-conflict QA data. It needs logits from two passes, with and without context, so it fits self-hosted models more than closed APIs.
Training. Longpre et al. trained on substituted examples and cut memorization to negligible levels while improving out-of-distribution F1 by 4 to 7%. They also found readers trained on less relevant passages learned to ignore the passage, and gold-passage training reduced that. Yoran et al. fine-tuned on a mix of relevant and irrelevant contexts and report that as few as 1,000 examples made the model robust to irrelevant context while keeping performance when the context was relevant. Self-RAG (Asai et al.) trains a model to critique retrieved passages and its own output with reflection tokens. Fine-tuning is the heavy option. Say when it is worth it: you own the model and prompt fixes have stalled on a measured set.
Verify and abstain. Anthropic's guide suggests having the model find a supporting quote for each claim after drafting and retract any claim it cannot support. Its Citations API returns cited text that is guaranteed to point at the provided documents. A valid pointer is still not proof that the claim matches the quote, so check numbers and names against the cited span. Ragas faithfulness scores the share of answer claims that can be inferred from the retrieved context. A faithful answer to a wrong chunk still scores 1.0, so pair faithfulness with answer correctness on your gold set. Joren et al. used a sufficient-context signal for selective generation and raised the fraction of correct answers among the questions the model chose to answer by 2 to 10% for Gemini, GPT, and Gemma.
What interviewers probe
Can you separate a retrieval miss from a generation failure with evidence? Do you know both directions of knowledge conflict, and that fixing one can worsen the other? Do you measure abstention on insufficient context instead of counting every answer? Do you know which fixes need logits or training, and which work on a hosted API? Do you know faithfulness is not correctness?
Sample interview answer
"First I check the logged prompt. If the gold chunk is not in it, it's a retrieval or ranking bug and I fix it there. If it is, I classify the failure. I run the question closed-book and open-book to see whether the prior is overriding correct context or a bad chunk is overriding a correct prior. ClashEval found models adopt wrong retrieved content over their own correct answer more than 60% of the time, so I track both prior bias and context bias. I do a counterfactual swap on a gold chunk to see if the model reads, and I move the gold chunk to the top because Lost in the Middle shows position matters. I label whether each context is sufficient, because the Sufficient Context paper showed strong models often answer instead of abstaining. Then I fix the layer: an 'I don't know' option, quote-first answering, and source-tagged documents above the question in the prompt. Fewer and reranked chunks in assembly. If we host the model, context-aware decoding or fine-tuning on counterfactual and noisy examples. A claim-level citation check with abstention at the end. I gate every change on the same frozen set with correctness, faithfulness, and abstention reported separately."
Follow-up questions to prepare
How would you tell that the model ignored the context from the context being stale? (Compare the chunk's source version with the corpus, then rerun closed-book.) Should a RAG system ever answer against its context? (Only with a flag, for example when the chunk is out of date, and that is a product decision.) Why can raising top-k make this worse? (More distractors and more middle positions.) How do you score an abstention? (As correct when the context is insufficient, as a miss when it was sufficient.) Where do multi-tenant permissions interact with this? (A filtered-out document looks like insufficient context. See Multi-tenant RAG isolation, https://aiinterviewquestion.com/blog/multi-tenant-rag-isolation.)
The close: prove the evidence was in the prompt, name the direction of the conflict, fix the layer that failed, and report correctness and abstention next to faithfulness. For the wider hallucination picture, see Why do LLMs hallucinate (https://aiinterviewquestion.com/blog/why-do-llms-hallucinate).
https://arxiv.org/abs/2109.05052 https://arxiv.org/abs/2305.13300 https://arxiv.org/abs/2404.10198 https://arxiv.org/abs/2303.11315 https://arxiv.org/abs/2305.14739 https://arxiv.org/abs/2307.03172 https://arxiv.org/abs/2302.00093 https://arxiv.org/abs/2310.01558 https://arxiv.org/abs/2411.06037 https://arxiv.org/abs/2310.11511 https://docs.anthropic.com/en/docs/test-and-evaluate/strengthen-guardrails/reduce-hallucinations https://docs.anthropic.com/en/docs/build-with-claude/prompt-engineering/claude-prompting-best-practices https://docs.anthropic.com/en/docs/build-with-claude/citations https://docs.ragas.io/en/stable/concepts/metrics/available_metrics/faithfulness/
Deep explanations with architecture diagrams for every question below.
Senior debugging interview question on Faithfulness within RAG.
Read full explanationSenior scenario interview question on Faithfulness within RAG.
Read full explanationSenior trade-off interview question on Context Precision within RAG.
Read full explanationSenior architecture interview question on Context Precision within RAG.
Read full explanationJunior conceptual interview question on Context Injection within RAG.
Read full explanationJunior system design interview question on Context Injection within RAG.
Read full explanationSenior trade-off interview question on Corrective RAG within RAG.
Read full explanationSenior debugging interview question on Corrective RAG within RAG.
Read full explanation