Executable programs and extensions

When local RAG answer is wrong: Search, chunk, model diagnosis order

If you change the model first, it is easy to miss the cause.

When an answer is wrong, first inspect the passages passed by retrieval. Check in order whether PDF text is missing, a chunk split a condition, or the required sentence was retrieved but the model missed it.

Requirements and key details
  • Open the retrieved passages that were sent to the model alongside the answer.
  • Compare the original with extracted text for tables and two-column PDFs.
  • Change one chunking, embedding, or filter setting at a time and rerun the same question.
  • Measure retrieval latency, time to first token, and generation time separately.

If an answer about the vacation policy is wrong, start by opening the retrieval results

Take the question, “Can an employee who has been here for six months take summer leave?” The answer screen alone does not reveal the cause, so first inspect in the UI or logs the passages retrieval sent to the model. If the top results do not contain the “one year of service” condition, do not switch the generation model. An answer model cannot recover a policy that was never retrieved.

If retrieval is empty or points to a different type of leave, compare the PDF extraction with the original page first. Check that headings and table rows have not been mixed up, and that conditions such as “six months or more” remain in the text. If the condition is in the extracted text but missing from retrieval, inspect chunk boundaries, embeddings, and document filters one at a time. If a retrieved passage clearly contains the condition but the answer is wrong, compare the prompt sent to the model with its answer.

RAG diagnosis that compares the original text, search evidence, and answers in order
Instead of changing the final answer, check first whether the basis for the correct answer has been found.

Compare the original PDF with extracted text first

If the vacation policy is a table or a two-column PDF, choose one page and place its visible text next to the extracted text. If “six months or more of service” is missing, or eligibility and exceptions from the table are mixed together, retrieval may be accurately finding incorrect source text. Change the PDF extraction tool or OCR settings, then extract that page again and check whether the missing sentence has been restored.

Before reindexing the entire document, first confirm that the text on the problem page is fixed. If the old result still appears after the correction, check whether the system is reading an earlier extraction or index. Move on to chunking and retrieval settings only after confirming that the original and extracted text convey the same meaning.

Comparison of messy sentences and organized extraction results from PDF original text
Two-column documents and tables are checked against the original text to ensure that the extraction order is not mixed up.

When evidence is not retrieved, change one setting at a time

If chunks are too short, the “six months of service” condition and whether leave is allowed may land in separate passages. If they are too long, several leave policies can be combined into one passage and unrelated sentences may be sent to the model. Mark the sentence that should support the answer, rerun the same question, and record whether it appears among the top retrieval results.

If evidence is missing, change only the chunk size and run the same question again. If it is still missing, change one item at a time in this order: overlap, embeddings, metadata filters, and reranker. If you change several settings at once, you will not know which one fixed retrieval even if the results change. Record improved retrieval ranking separately from improved answer accuracy.

Verify the new index before switching over after a document update

If the old answer continues after you revise the policy PDF, check the file hash, index creation time, index version, and cache key rather than relying on the filename. The contents may have changed while the name stayed the same. Use a question to verify that the revised document appears in the extracted text and that the updated clause appears in results from the new index.

If you delete the active index before rebuilding it, retrieval may be unavailable until validation is complete. Build a separate version, check it with representative questions, and then switch over. If the new version has a problem, you can return to the previous index. Verify the document hash, index version, and retrieved passages and answers for representative questions.

Reviewing a new index version while preserving the old index
The new index must be verified separately and then converted so that you can return to it immediately if a problem occurs.

If retrieval is correct but the answer is wrong, inspect generation

Suppose the retrieved passage says, “Employees may request summer leave after six months of service,” but the model answers, “You must have worked for at least one year.” First check that the exact passage was included in the actual prompt. If it was, instruct the model to answer only from the evidence and to say it does not know when the condition is absent. If it is still wrong with the same evidence, compare a different model or quantization1.

Break down slow responses as well. Record the time from sending the question until retrieval results appear, from the end of retrieval until the first token2, and from the first token until the answer is complete. If the first interval is long, inspect the index and filters; if the second is long, check evidence length and prefill3; if the last is long, check token generation speed. If slowness occurs only with more users, compare queueing and memory use when one person and four people ask the same question.

Terminology notes

  1. Quantization — Representing model values with fewer bits. Memory use, accuracy, or execution speed may change; the effects depend on the format and implementation.

    Back to the text
  2. Token — A unit into which a model divides input or output for processing. One token does not equal one character or a fixed duration.

    Back to the text
  3. Prefill — The stage where an LLM reads the input prompt and computes representations for its tokens. Longer prompts contain more tokens to process.

    Back to the text