1Purpose first
Two places a RAG system can fail.
A retrieval-augmented system can fail in two different places, and a wrong answer does not tell you which. Evaluation is how you find out. It is not a checkbox after the build. It tells you which half of the system actually broke.
2Four ideas worth keeping
From the course, in one table.
Four ideas from the course that are worth keeping.
| Idea | What it means | Why it matters |
|---|---|---|
| Faithfulness and relevance are separate | An answer can be fully grounded in the retrieved text and still not answer the question | Check both, or a faithful but off-topic answer passes |
| Precision and recall trade off | Retrieve tighter and you may miss things. Cast wider and you let in noise | Track both, or you only see half the picture |
| Index dimensions must match the embedding model | A mismatch fails, and often late | Cheap to check early, expensive to debug later |
| Measure, then change | Fix a test set, change one thing, compare scores | Gut feel is not a result |
3Precision and recall
Retrieval scores on a tiny test.
For retrieval, the two classic numbers are precision (of what I fetched, how much was useful?) and recall (of what was useful, how much did I fetch?). Here they are on a tiny labelled test.
# a tiny labelled test: for each question, which chunks are truly relevant,
# and which chunks the retriever returned (top 3)
tests = [
{"q": "Can I return shoes?", "relevant": {"c1", "c2"}, "returned": ["c1", "c9", "c7"]},
{"q": "Who pays shipping?", "relevant": {"c3"}, "returned": ["c3", "c2", "c5"]},
]
def precision(relevant, returned): # of what I returned, how much was useful?
return len([c for c in returned if c in relevant]) / len(returned)
def recall(relevant, returned): # of what was useful, how much did I return?
return len([c for c in returned if c in relevant]) / len(relevant)
for t in tests:
p, r = precision(t["relevant"], t["returned"]), recall(t["relevant"], t["returned"])
print(f"{t['q']:<22} precision {p:.2f} recall {r:.2f}")
avg_p = sum(precision(t["relevant"], t["returned"]) for t in tests) / len(tests)
avg_r = sum(recall(t["relevant"], t["returned"]) for t in tests) / len(tests)
print(f"average precision {avg_p:.2f} recall {avg_r:.2f}")Can I return shoes? precision 0.33 recall 0.50 Who pays shipping? precision 0.33 recall 1.00 average precision 0.33 recall 0.75
Reading it: the first question got one right chunk out of three fetched (precision 0.33) but only one of its two needed chunks (recall 0.50). The second found its single needed chunk (recall 1.00) while fetching two others that did not help. Which number to push depends on the system. Raising the number of chunks fetched tends to raise recall and lower precision.
4Generation scores
Faithfulness and relevance, and how to read them together.
Retrieval scores tell you if the right text arrived. Generation scores tell you what the model did with it.
| Score | Question it answers | Typical name |
|---|---|---|
| Faithfulness / groundedness | Is every claim in the answer supported by the retrieved text? | Ragas: faithfulness. Triad: groundedness |
| Answer relevance | Does the answer address what was asked? | Ragas: response relevancy. Triad: answer relevance |
| Context precision / relevance | Is the retrieved text on topic? | Ragas: context precision. Triad: context relevance |
| Context recall | Did retrieval capture what was needed? | Ragas: context recall |
Two vendors, two sets of names, mostly the same ideas. Knowing both means you can translate between tools.
5Dimensions must match
A cheap check that saves a bad day.
An index is created for vectors of a fixed length. If the embedding model produces a different length, the write or search fails, sometimes only later in the pipeline. A one-line check at the start of a build is much cheaper than finding it in production.
INDEX_DIMENSION = 1536 # fixed when the index was created
def check_embedding(vector, model_name):
if len(vector) != INDEX_DIMENSION:
raise ValueError(f"{model_name} gives {len(vector)} numbers, index expects {INDEX_DIMENSION}")
return True
print(check_embedding([0.0] * 1536, "model-A"))
try:
check_embedding([0.0] * 1024, "model-B")
except ValueError as e:
print("caught early:", e)True caught early: model-B gives 1024 numbers, index expects 1536
The same idea applies whenever you change embedding models: the index has to be rebuilt, because vectors from different models are not comparable.
6A domain-tuned embedder?
An idea to test, not a rule.
One more point from the course, which I have not tested myself: if you train a model on your own domain, you can read an embedding from a different layer of that same network and get a domain-tuned embedder as a by-product. Treat it as an idea to explore, and measure it against a general embedding model on your own questions before relying on it.
7The loop
Baseline, change, measure, repeat.
Keep the question set small enough to read through by hand, and varied enough to include the awkward cases. Change one variable at a time, or you will not know what helped.
8Check yourself
Short questions and one line to remember.
Why measure faithfulness and relevance separately?
An answer can be grounded in the retrieved text and still miss the question. One score hides that.
What happens to precision when you fetch more chunks?
It usually falls, while recall rises. They trade off, so track both.
Why check embedding dimensions early?
A mismatch fails, and if you find it late it is costly to trace. Check at build time.
If answers are wrong but retrieved context looks right, what do you fix?
The prompt or the model, since retrieval did its job.