1Purpose first
What RAG is and where quality comes from.
A language model only knows what it was trained on and what you put in the prompt. Retrieval-augmented generation (RAG) fills the gap: find the passages of your own documents that matter for the question, put them in the prompt, and have the model answer from them.
Poor answers usually trace back to retrieval, not the model. Retrieval has three layers, and each fixes a different failure: chunking (how documents are cut), search (how pieces are found) and reranking (how the finalists are ordered).
2The pipeline
Three layers, in order.
3Small versus large
Why size has no default.
There is no best chunk size, because two things pull in opposite directions.
| Small chunks | Large chunks | |
|---|---|---|
| Precision | High: a hit is on-topic | Lower: one chunk mixes topics |
| Context | Low: the answer may be split across chunks | High: the surrounding detail is there |
| Cost | More chunks to store and search | Fewer chunks, more tokens per answer |
| Typical failure | "Not refunded" retrieved without saying what is not refunded | Right chunk found, buried among noise |
doc = ["Refunds are allowed within 30 days.", "Items must be unused.",
"Shipping costs are not refunded.", "Digital goods cannot be refunded.",
"Contact support to start a refund."]
def chunk(sents, n):
return [" ".join(sents[i:i + n]) for i in range(0, len(sents), n)]
def window(sents, i, w=1): # sentence-window: the hit plus its neighbours
return " ".join(sents[max(0, i - w): i + w + 1])
print("1 sentence per chunk :", len(chunk(doc, 1)), "chunks")
print("3 sentences per chunk:", len(chunk(doc, 3)), "chunks")
print("window around #3 :", window(doc, 2))1 sentence per chunk : 5 chunks 3 sentences per chunk: 2 chunks window around #3 : Items must be unused. Shipping costs are not refunded. Digital goods cannot be refunded.
One sentence per chunk gives five precise pieces, but "Shipping costs are not refunded" no longer says refunded from what. Three per chunk keeps context, but a question about one rule pulls in a mixed bag. The last line previews a fix covered in Advanced retrieval.
4Embeddings and vector stores
Where the pieces are kept.
To search by meaning, each chunk is turned into an embedding, a list of numbers where similar meanings sit close together (see How a transformer works). The embeddings go into a vector store, which finds the closest ones to a question quickly.
| Kind | Examples | Trade-off |
|---|---|---|
| Runs inside your app or on your machine | Chroma, the vector search in your existing database | Simple and free to start; you look after it |
| Managed service | Pinecone | Less to operate; you pay and depend on a vendor |
5Search: keyword, meaning, hybrid
Two ways to find, merged.
Two ways to find candidates, and each misses what the other catches.
| Keyword search | Meaning (vector) search | |
|---|---|---|
| Finds | Exact words and codes | Paraphrases and related ideas |
| Misses | "package arrives" when the text says "shipping takes" | Exact identifiers such as RF-30 |
| Example use | Product codes, names, error messages | Natural-language questions |
Hybrid search runs both and merges the two ranked lists. A common merge is reciprocal rank fusion: each chunk scores higher the nearer the top it sits in either list. The toy below shows each method rescuing the other's miss.
docs = [
"Shipping takes 3 to 5 working days within India.",
"Shipping costs are not refunded.",
"Refunds are issued within 30 days of purchase.",
"Code RF-30 applies to refund requests.",
]
# a toy "meaning" map: words that mean the same thing share a concept
concepts = {"shipping": "delivery", "package": "delivery", "arrives": "delivery", "arrive": "delivery",
"days": "time", "long": "time", "takes": "time", "refunded": "refund", "refunds": "refund",
"refund": "refund"}
def words(t): return [w.strip(".,?").lower() for w in t.split()]
def keyword(q, d): return len(set(words(q)) & set(words(d))) # layer 2a: exact words
def meaning(q, d): # layer 2b: shared concepts
cq = {concepts[w] for w in words(q) if w in concepts}
cd = {concepts[w] for w in words(d) if w in concepts}
return len(cq & cd)
def rank(score, q):
order = sorted(range(len(docs)), key=lambda i: -score(q, docs[i]))
return [i for i in order if score(q, docs[i]) > 0]
def fuse(lists, k=60): # reciprocal rank fusion
s = {}
for lst in lists:
for pos, i in enumerate(lst):
s[i] = s.get(i, 0) + 1 / (k + pos + 1)
return sorted(s, key=lambda i: -s[i])
for q in ["how long until my package arrives", "what is code RF-30"]:
kw, mn = rank(keyword, q), rank(meaning, q)
print(q)
print(" keyword only:", kw)
print(" meaning only:", mn)
print(" fused :", fuse([kw, mn]))how long until my package arrives keyword only: [] meaning only: [0, 1, 2] fused : [0, 1, 2] what is code RF-30 keyword only: [3] meaning only: [] fused : [3]
The first question is a paraphrase, where keywords find nothing. The second is an exact code, where meaning finds nothing. Hybrid gets both. Many setups also weight one side more heavily, with meaning a little ahead being a common start. Tune the weights on your own questions.
6Reranking
A sharper order for the shortlist.
Search is built for speed over many chunks, so its ordering is rough. A reranker is slower but sharper. It reads the question and each candidate together and rescores them, and you keep only the best few.
# layer 3: a reranker reads the question and each candidate TOGETHER and rescores.
# Real rerankers are trained models; this toy just rewards covering more of the question.
question = "how many days do refunds take"
candidates = [ # what the search layer returned, in its order
"Shipping costs are not refunded.",
"Refunds are issued within 30 days of purchase.",
"Our office is open on weekdays.",
]
stem = lambda w: w.strip(".,?").lower().rstrip("s")
q = {stem(w) for w in question.split()} - {"how", "many", "do"}
def score(c):
return len(q & {stem(w) for w in c.split()}) / len(q)
print("before:", candidates[0])
ranked = sorted(candidates, key=score, reverse=True)
print("after :", ranked[0])
for c in ranked: print(round(score(c), 2), c)before: Shipping costs are not refunded. after : Refunds are issued within 30 days of purchase. 0.67 Refunds are issued within 30 days of purchase. 0.0 Shipping costs are not refunded. 0.0 Our office is open on weekdays.
The toy rewards covering more of the question. Real rerankers are trained models. For example Cohere's Rerank takes a query, a list of documents and a top_n, and returns them in relevance order. Its documentation currently recommends rerank-v4.0-pro, with a faster rerank-v4.0-fast for latency-sensitive uses.
7Diagnose by symptom
Which layer to touch first.
| Symptom | Likely layer | First thing to try |
|---|---|---|
| Answer is half right; the rest was in the next paragraph | Chunking | Larger chunks or sentence-window retrieval |
| Right chunk exists but never comes back | Search | Add keyword search; check embedding model |
| Right chunk is retrieved but ranked low | Reranking | Add a reranker over the top candidates |
Change one layer at a time and measure each change on a fixed set of real questions (see Evaluating RAG).
8Check yourself
Short questions and one line to remember.
Why is there no universal chunk size?
Small chunks are precise but lose context. Large ones keep context but mix topics. The right size depends on your documents and questions.
Why use hybrid search?
Keyword search catches exact terms that meaning search misses, and meaning search catches paraphrases that keywords miss.
What does a reranker add?
A sharper second ordering of a short list, by reading the question and each candidate together.