← Home · Tutorial 14 · Retrieval (RAG)

Evaluating RAG

Precision, recall, faithfulness and relevance, and how to tell which half of the system broke.

Last reviewed 5 October 2026 · Written with AI assistance and reviewed by me

Where I learned this: my notes from finishing Activeloop's RAG course, which stressed evaluation as part of the build. I added the metric definitions from the Ragas documentation and the "RAG triad" from DeepLearning.AI's Building and Evaluating Advanced RAG. The examples and code are my own.
Who this is for: anyone whose retrieval-based chatbot gives answers they cannot yet measure. You will learn precision and recall, generation scores, and the loop of changing one thing and re-measuring.

1Purpose first

Two places a RAG system can fail.

A retrieval-augmented system can fail in two different places, and a wrong answer does not tell you which. Evaluation is how you find out. It is not a checkbox after the build. It tells you which half of the system actually broke.

Two places it can break
Question
→
Retrievaldid we fetch the right text?
→
Generationdid the model use it well?
→
Answer

2Four ideas worth keeping

From the course, in one table.

Four ideas from the course that are worth keeping.

IdeaWhat it meansWhy it matters
Faithfulness and relevance are separateAn answer can be fully grounded in the retrieved text and still not answer the questionCheck both, or a faithful but off-topic answer passes
Precision and recall trade offRetrieve tighter and you may miss things. Cast wider and you let in noiseTrack both, or you only see half the picture
Index dimensions must match the embedding modelA mismatch fails, and often lateCheap to check early, expensive to debug later
Measure, then changeFix a test set, change one thing, compare scoresGut feel is not a result

3Precision and recall

Retrieval scores on a tiny test.

For retrieval, the two classic numbers are precision (of what I fetched, how much was useful?) and recall (of what was useful, how much did I fetch?). Here they are on a tiny labelled test.

Python
# a tiny labelled test: for each question, which chunks are truly relevant,
# and which chunks the retriever returned (top 3)
tests = [
    {"q": "Can I return shoes?", "relevant": {"c1", "c2"}, "returned": ["c1", "c9", "c7"]},
    {"q": "Who pays shipping?",  "relevant": {"c3"},       "returned": ["c3", "c2", "c5"]},
]

def precision(relevant, returned):            # of what I returned, how much was useful?
    return len([c for c in returned if c in relevant]) / len(returned)

def recall(relevant, returned):               # of what was useful, how much did I return?
    return len([c for c in returned if c in relevant]) / len(relevant)

for t in tests:
    p, r = precision(t["relevant"], t["returned"]), recall(t["relevant"], t["returned"])
    print(f"{t['q']:<22} precision {p:.2f}  recall {r:.2f}")

avg_p = sum(precision(t["relevant"], t["returned"]) for t in tests) / len(tests)
avg_r = sum(recall(t["relevant"], t["returned"]) for t in tests) / len(tests)
print(f"average                precision {avg_p:.2f}  recall {avg_r:.2f}")
Output (run to check)
Can I return shoes?    precision 0.33  recall 0.50
Who pays shipping?     precision 0.33  recall 1.00
average                precision 0.33  recall 0.75

Reading it: the first question got one right chunk out of three fetched (precision 0.33) but only one of its two needed chunks (recall 0.50). The second found its single needed chunk (recall 1.00) while fetching two others that did not help. Which number to push depends on the system. Raising the number of chunks fetched tends to raise recall and lower precision.

4Generation scores

Faithfulness and relevance, and how to read them together.

Retrieval scores tell you if the right text arrived. Generation scores tell you what the model did with it.

ScoreQuestion it answersTypical name
Faithfulness / groundednessIs every claim in the answer supported by the retrieved text?Ragas: faithfulness. Triad: groundedness
Answer relevanceDoes the answer address what was asked?Ragas: response relevancy. Triad: answer relevance
Context precision / relevanceIs the retrieved text on topic?Ragas: context precision. Triad: context relevance
Context recallDid retrieval capture what was needed?Ragas: context recall
Reading the pattern: poor context scores with decent answers usually means retrieval needs work (chunking, search, reranking, see tutorial 12). Good context scores with poor answers points at the prompt or the model. Without separate scores you would tune the wrong half.

Two vendors, two sets of names, mostly the same ideas. Knowing both means you can translate between tools.

5Dimensions must match

A cheap check that saves a bad day.

An index is created for vectors of a fixed length. If the embedding model produces a different length, the write or search fails, sometimes only later in the pipeline. A one-line check at the start of a build is much cheaper than finding it in production.

Python
INDEX_DIMENSION = 1536                         # fixed when the index was created

def check_embedding(vector, model_name):
    if len(vector) != INDEX_DIMENSION:
        raise ValueError(f"{model_name} gives {len(vector)} numbers, index expects {INDEX_DIMENSION}")
    return True

print(check_embedding([0.0] * 1536, "model-A"))
try:
    check_embedding([0.0] * 1024, "model-B")
except ValueError as e:
    print("caught early:", e)
Output (run to check)
True
caught early: model-B gives 1024 numbers, index expects 1536

The same idea applies whenever you change embedding models: the index has to be rebuilt, because vectors from different models are not comparable.

6A domain-tuned embedder?

An idea to test, not a rule.

One more point from the course, which I have not tested myself: if you train a model on your own domain, you can read an embedding from a different layer of that same network and get a domain-tuned embedder as a by-product. Treat it as an idea to explore, and measure it against a general embedding model on your own questions before relying on it.

7The loop

Baseline, change, measure, repeat.

The evaluation loop
Fixed question setwith known good answers
→
Baseline scoresretrieval and generation
→
Change one thingchunking, search, prompt
→
Re-measurecompare, keep or revert
↺

Keep the question set small enough to read through by hand, and varied enough to include the awkward cases. Change one variable at a time, or you will not know what helped.

8Check yourself

Short questions and one line to remember.

Why measure faithfulness and relevance separately?

An answer can be grounded in the retrieved text and still miss the question. One score hides that.

What happens to precision when you fetch more chunks?

It usually falls, while recall rises. They trade off, so track both.

Why check embedding dimensions early?

A mismatch fails, and if you find it late it is costly to trace. Check at build time.

If answers are wrong but retrieved context looks right, what do you fix?

The prompt or the model, since retrieval did its job.

One line to remember: "Evaluation tells you which half of the system broke, so measure retrieval and generation separately."
Try it yourself: write ten questions with the chunks that should answer them, run your retriever, and compute precision and recall before changing anything.
Notes based on Activeloop’s RAG course, with metric definitions from the Ragas documentation and the RAG triad from DeepLearning.AI’s Building and Evaluating Advanced RAG. Independent notes, not affiliated with or endorsed by any company or course named here. Examples and wording are my own.