Executable programs and extensions
Local LLM·RAG quality evaluation: Verifying changed settings with DeepEval
We need to check separately whether the quick settings have improved business response.
Keep the work questions and supporting evidence fixed, then use DeepEval to score retrieved context and answers separately. You can designate Ollama as the judge model, but review the source and response-time and memory records alongside the scores.
Write down 20 vacation-policy questions and their answers first
As an example, we will evaluate a RAG1 system that searches an internal vacation-policy PDF. Prepare 20 questions, including ones answered by a single sentence, ones requiring an exception in a table, and ones the document cannot answer. For each row, record the question, expected answer, and the supporting sentence that must be retrieved, or “no evidence in the document.” Base the set on real work questions, but first check whether personal information or confidential source documents may be sent to the evaluation service.
Save a baseline run, then change only one of the model, quantization2, or chunk settings. Use the same questions and expected answers for both configurations so differences in the question set are not mistaken for performance changes. If the document or expected answers change, version the question set and record that change as well.

Score retrieved evidence and answers separately
To distinguish retrieval failures from answer failures, include the input, expected answer, model answer, and retrieved evidence in the evaluation data. DeepEval's ContextualRecallMetric checks whether the retrieved context contains information needed for the expected answer. FaithfulnessMetric checks whether the generated answer stays within that context. AnswerRelevancyMetric can check whether the answer directly addresses the question. These metrics measure different things, so review each question's scores and the reasons given side by side.
Do not use RAG scores alone to determine whether a policy has been violated. The judge model is an aid for comparing sentence meaning, not the answer key itself. For questions where omitting a required condition such as “six months or more” must count as a failure, have a person check the actual retrieved sentence and answer. DeepEval's LLM3-based metrics call a separate judge model, so running the application locally does not automatically make evaluation local too.

Set a local judge model and run the first test
DeepEval's official documentation explains how to use Ollama as a local judge model. Install Ollama and start its server, then install DeepEval in a new Python virtual environment4 for the project. ollama pull deepseek-r1:1.5b downloads the example judge model, while deepeval set-ollama --model=deepseek-r1:1.5b selects it as DeepEval's default evaluation model. If you already have a model, use the name shown by ollama list. The goal is to evaluate your application's answer quality, so replace the example answer and retrieved passage in the test code below with results from your actual RAG run.
The following example tests one answer based on an eligibility condition in the vacation policy. Save the file as test_rag.py and run deepeval test run test_rag.py to confirm that the Ollama judge is connected. Install and download the model in an environment with network access. If you cannot obtain the model or verify the judge connection in an air-gapped environment, you have not confirmed that the evaluation ran. The metric thresholds are examples; set them based on your organization's required questions and human review.
If you cannot prepare an Ollama model and run the commands above, do not treat this as a completed local evaluation. Preparing the test code does not produce results until the judge-model connection is verified. Also, signing in to or syncing results with an external service such as Confident AI can send run records off the device. In an air-gapped environment, separately verify local result storage and that network access is blocked.
python -m venv .venv
source .venv/bin/activate
pip install -U deepeval
ollama pull deepseek-r1:1.5b
deepeval set-ollama --model=deepseek-r1:1.5bfrom deepeval import assert_test
from deepeval.metrics import (
AnswerRelevancyMetric,
ContextualRecallMetric,
FaithfulnessMetric,
)
from deepeval.test_case import LLMTestCase
def test_vacation_policy_answer():
test_case = LLMTestCase(
input="입사 6개월 차 직원도 여름휴가를 쓸 수 있나요?",
actual_output="입사 후 6개월 이상이면 여름휴가를 신청할 수 있습니다.",
expected_output="입사 후 6개월 이상이면 여름휴가를 신청할 수 있습니다.",
retrieval_context=[
"여름휴가는 입사 후 6개월 이상 근무한 직원이 신청할 수 있습니다."
],
)
assert_test(
test_case,
[
ContextualRecallMetric(threshold=0.5),
FaithfulnessMetric(threshold=0.5),
AnswerRelevancyMetric(threshold=0.5),
],
)Set acceptance criteria before comparing speed and memory
For each run, record the answer model and quantization, chunk size, embedding model, judge-model name and version, time to first token5, generation speed, retrieval time, and peak memory. Evaluation scores alone do not tell you how long the faster configuration makes a real user wait; speed alone can hide a missed critical condition.
Do not choose based on one overall average. First set a failure-count or pass-rate criterion for questions that must be answered correctly, such as leave eligibility. Then compare speed and memory only among configurations that meet it. For example, you may reject a configuration that presents “no evidence” as a supported answer to a required question, even if its average score is high. If score differences are small or the judge's reasoning seems off, have a person inspect the source, retrieved passage, and answer before comparing the next setting.

Terminology notes
Retrieval-Augmented Generation — A method that retrieves material relevant to a query, adds it to the model input, and generates an answer. Search scope and source quality depend on the implementation.
Back to the textQuantization — Representing model values with fewer bits. Memory use, accuracy, or execution speed may change; the effects depend on the format and implementation.
Back to the textLarge language model — A language model trained on large text datasets to process and generate text. Capabilities and supported inputs vary by model.
Back to the textPython virtual environment — An isolated space for installing Python packages per project. It helps reduce version conflicts and is not a virtual machine.
Back to the textTime to First Token — The time from sending a request until its first output token arrives. It is separate from the generation rate of later tokens.
Back to the text