← Home · Tutorial 12 · Retrieval (RAG)

Retrieval in three layers

Chunking, search and reranking: what each layer fixes, and how to tell which one to change.

Last reviewed 5 October 2026 · Written with AI assistance and reviewed by me

Where this comes from: my notes on Activeloop's RAG course (chunking) and on retrieval work I did hands-on, checked against Cohere's Rerank documentation. The explanations and toy code are mine. The code uses small hand-made examples to show the idea, not real embeddings or a real reranker.
Who this is for: anyone building a chatbot or search over their own documents. You will learn the three layers that decide answer quality (chunking, search, reranking) and what each one fixes.

1Purpose first

What RAG is and where quality comes from.

A language model only knows what it was trained on and what you put in the prompt. Retrieval-augmented generation (RAG) fills the gap: find the passages of your own documents that matter for the question, put them in the prompt, and have the model answer from them.

Poor answers usually trace back to retrieval, not the model. Retrieval has three layers, and each fixes a different failure: chunking (how documents are cut), search (how pieces are found) and reranking (how the finalists are ordered).

2The pipeline

Three layers, in order.

Three layers
Documentspolicies, manuals
→
1. Chunkcut into pieces
→
2. Searchfind candidates
→
3. Rerankorder the finalists
→
Answermodel reads the top few

3Small versus large

Why size has no default.

There is no best chunk size, because two things pull in opposite directions.

Small chunksLarge chunks
PrecisionHigh: a hit is on-topicLower: one chunk mixes topics
ContextLow: the answer may be split across chunksHigh: the surrounding detail is there
CostMore chunks to store and searchFewer chunks, more tokens per answer
Typical failure"Not refunded" retrieved without saying what is not refundedRight chunk found, buried among noise
Python
doc = ["Refunds are allowed within 30 days.", "Items must be unused.",
       "Shipping costs are not refunded.", "Digital goods cannot be refunded.",
       "Contact support to start a refund."]

def chunk(sents, n):
    return [" ".join(sents[i:i + n]) for i in range(0, len(sents), n)]

def window(sents, i, w=1):                   # sentence-window: the hit plus its neighbours
    return " ".join(sents[max(0, i - w): i + w + 1])

print("1 sentence per chunk :", len(chunk(doc, 1)), "chunks")
print("3 sentences per chunk:", len(chunk(doc, 3)), "chunks")
print("window around #3     :", window(doc, 2))
Output (run to check)
1 sentence per chunk : 5 chunks
3 sentences per chunk: 2 chunks
window around #3     : Items must be unused. Shipping costs are not refunded. Digital goods cannot be refunded.

One sentence per chunk gives five precise pieces, but "Shipping costs are not refunded" no longer says refunded from what. Three per chunk keeps context, but a question about one rule pulls in a mixed bag. The last line previews a fix covered in Advanced retrieval.

Before anything clever: cut on the document's own structure. A FAQ should be one question and answer per chunk. Fixed-size cutting can glue unrelated topics into one chunk, and no later layer can fully repair that.
The trade-off in one line: small chunks are easy to find, large chunks are easy to understand.

4Embeddings and vector stores

Where the pieces are kept.

To search by meaning, each chunk is turned into an embedding, a list of numbers where similar meanings sit close together (see How a transformer works). The embeddings go into a vector store, which finds the closest ones to a question quickly.

KindExamplesTrade-off
Runs inside your app or on your machineChroma, the vector search in your existing databaseSimple and free to start; you look after it
Managed servicePineconeLess to operate; you pay and depend on a vendor
One check that saves a day: the index dimension must match your embedding model's output size. Switch embedding models and an old index will reject the new vectors.

5Search: keyword, meaning, hybrid

Two ways to find, merged.

Two ways to find candidates, and each misses what the other catches.

Keyword searchMeaning (vector) search
FindsExact words and codesParaphrases and related ideas
Misses"package arrives" when the text says "shipping takes"Exact identifiers such as RF-30
Example useProduct codes, names, error messagesNatural-language questions

Hybrid search runs both and merges the two ranked lists. A common merge is reciprocal rank fusion: each chunk scores higher the nearer the top it sits in either list. The toy below shows each method rescuing the other's miss.

Python
docs = [
    "Shipping takes 3 to 5 working days within India.",
    "Shipping costs are not refunded.",
    "Refunds are issued within 30 days of purchase.",
    "Code RF-30 applies to refund requests.",
]
# a toy "meaning" map: words that mean the same thing share a concept
concepts = {"shipping": "delivery", "package": "delivery", "arrives": "delivery", "arrive": "delivery",
            "days": "time", "long": "time", "takes": "time", "refunded": "refund", "refunds": "refund",
            "refund": "refund"}

def words(t): return [w.strip(".,?").lower() for w in t.split()]
def keyword(q, d): return len(set(words(q)) & set(words(d)))                 # layer 2a: exact words
def meaning(q, d):                                                           # layer 2b: shared concepts
    cq = {concepts[w] for w in words(q) if w in concepts}
    cd = {concepts[w] for w in words(d) if w in concepts}
    return len(cq & cd)

def rank(score, q):
    order = sorted(range(len(docs)), key=lambda i: -score(q, docs[i]))
    return [i for i in order if score(q, docs[i]) > 0]

def fuse(lists, k=60):                                                       # reciprocal rank fusion
    s = {}
    for lst in lists:
        for pos, i in enumerate(lst):
            s[i] = s.get(i, 0) + 1 / (k + pos + 1)
    return sorted(s, key=lambda i: -s[i])

for q in ["how long until my package arrives", "what is code RF-30"]:
    kw, mn = rank(keyword, q), rank(meaning, q)
    print(q)
    print("  keyword only:", kw)
    print("  meaning only:", mn)
    print("  fused       :", fuse([kw, mn]))
Output (run to check)
how long until my package arrives
  keyword only: []
  meaning only: [0, 1, 2]
  fused       : [0, 1, 2]
what is code RF-30
  keyword only: [3]
  meaning only: []
  fused       : [3]

The first question is a paraphrase, where keywords find nothing. The second is an exact code, where meaning finds nothing. Hybrid gets both. Many setups also weight one side more heavily, with meaning a little ahead being a common start. Tune the weights on your own questions.

6Reranking

A sharper order for the shortlist.

Search is built for speed over many chunks, so its ordering is rough. A reranker is slower but sharper. It reads the question and each candidate together and rescores them, and you keep only the best few.

Python
# layer 3: a reranker reads the question and each candidate TOGETHER and rescores.
# Real rerankers are trained models; this toy just rewards covering more of the question.
question = "how many days do refunds take"
candidates = [                      # what the search layer returned, in its order
    "Shipping costs are not refunded.",
    "Refunds are issued within 30 days of purchase.",
    "Our office is open on weekdays.",
]
stem = lambda w: w.strip(".,?").lower().rstrip("s")
q = {stem(w) for w in question.split()} - {"how", "many", "do"}

def score(c):
    return len(q & {stem(w) for w in c.split()}) / len(q)

print("before:", candidates[0])
ranked = sorted(candidates, key=score, reverse=True)
print("after :", ranked[0])
for c in ranked: print(round(score(c), 2), c)
Output (run to check)
before: Shipping costs are not refunded.
after : Refunds are issued within 30 days of purchase.
0.67 Refunds are issued within 30 days of purchase.
0.0 Shipping costs are not refunded.
0.0 Our office is open on weekdays.

The toy rewards covering more of the question. Real rerankers are trained models. For example Cohere's Rerank takes a query, a list of documents and a top_n, and returns them in relevance order. Its documentation currently recommends rerank-v4.0-pro, with a faster rerank-v4.0-fast for latency-sensitive uses.

Rerank a shortlist, not everything. It costs time and money per candidate. Search pulls maybe twenty to fifty, the reranker keeps three to five.

7Diagnose by symptom

Which layer to touch first.

SymptomLikely layerFirst thing to try
Answer is half right; the rest was in the next paragraphChunkingLarger chunks or sentence-window retrieval
Right chunk exists but never comes backSearchAdd keyword search; check embedding model
Right chunk is retrieved but ranked lowRerankingAdd a reranker over the top candidates

Change one layer at a time and measure each change on a fixed set of real questions (see Evaluating RAG).

8Check yourself

Short questions and one line to remember.

Why is there no universal chunk size?

Small chunks are precise but lose context. Large ones keep context but mix topics. The right size depends on your documents and questions.

Why use hybrid search?

Keyword search catches exact terms that meaning search misses, and meaning search catches paraphrases that keywords miss.

What does a reranker add?

A sharper second ordering of a short list, by reading the question and each candidate together.

Try it: take one document you know well and write five questions it can answer. Cut it into one-sentence and three-sentence chunks, and note which version retrieves the right text for each question.
One line to remember: "Chunking decides what can be found, search finds candidates, and reranking decides what the model reads."
Notes based on Activeloop’s RAG course (chunking) and hands-on retrieval experiments, with reranking details from Cohere’s Rerank documentation. Model names and vendors change, so check current docs. Independent notes, not affiliated with or endorsed by any company or course named here. Examples and wording are my own.