← Home · Tutorial 13 · Retrieval (RAG)

Advanced retrieval

Sentence-window and auto-merging retrieval: search on small pieces, return bigger ones.

Last reviewed 5 October 2026 · Written with AI assistance and reviewed by me

Where this comes from: my notes on DeepLearning.AI's short course Building and Evaluating Advanced RAG, taught by Jerry Liu (LlamaIndex) and Anupam Datta (then of TruEra). The explanations and code are mine, and the code is a toy version of the idea, not their implementation.
Who this is for: people who already know chunking and search (see Retrieval in three layers) and want better context for the model. You will learn sentence-window retrieval, auto-merging retrieval, and when each helps.

1Purpose first

Escape the chunk-size choice.

Plain retrieval forces one choice of chunk size. Small chunks match well but give the model little to work with. Big chunks give context but match poorly. The techniques here refuse the choice: search on small pieces, then hand the model something bigger.

2Sentence-window retrieval

Hit plus neighbours.

Index single sentences so matching is precise. When a sentence matches, return it together with the sentences around it.

Python
doc = ["Refunds are allowed within 30 days.", "Items must be unused.",
       "Shipping costs are not refunded.", "Digital goods cannot be refunded.",
       "Contact support to start a refund."]

def window(sents, i, w=1):
    # return the matched sentence plus w sentences either side
    return " ".join(sents[max(0, i - w): i + w + 1])

hit = 2                                   # say the search matched this sentence
print("matched  :", doc[hit])
print("returned :", window(doc, hit))
print("wider    :", window(doc, hit, w=2))
Output (run to check)
matched  : Shipping costs are not refunded.
returned : Items must be unused. Shipping costs are not refunded. Digital goods cannot be refunded.
wider    : Refunds are allowed within 30 days. Items must be unused. Shipping costs are not refunded. Digital goods cannot be refunded. Contact support to start a refund.

The window is fixed, so it only helps when the answer sits in the neighbouring lines. If the relevant sentences are far apart, it will not connect them.

3Auto-merging retrieval

Merge children into the parent.

Build a hierarchy: small chunks (children) inside larger sections (parents). Index and match the small ones. If enough children of the same parent are retrieved, swap them for the parent. The more of a section that matches, the more of it you get.

Python
# small chunks (children) sit inside larger sections (parents)
parents = {
    "refunds": ["Refunds are allowed within 30 days.", "Items must be unused.", "Refunds go to the original payment method."],
    "shipping": ["Shipping takes 3 to 5 days.", "Tracking is sent by email.", "Remote areas can take longer."],
}
child_of = {c: p for p, kids in parents.items() for c in kids}

def auto_merge(retrieved, threshold=0.5):
    out, by_parent = [], {}
    for c in retrieved:
        by_parent.setdefault(child_of[c], []).append(c)
    for p, hits in by_parent.items():
        if len(hits) / len(parents[p]) >= threshold:
            out.append("[merged] " + " ".join(parents[p]))      # enough siblings matched: return the parent
        else:
            out.extend(hits)                                    # otherwise keep the small hits
    return out

retrieved = ["Refunds are allowed within 30 days.", "Items must be unused.", "Tracking is sent by email."]
for line in auto_merge(retrieved): print(line)
Output (run to check)
[merged] Refunds are allowed within 30 days. Items must be unused. Refunds go to the original payment method.
Tracking is sent by email.

Two of the three refund sentences matched, which is at least half, so the whole refund section was returned. Only one shipping sentence matched, so it stayed small. The threshold is the setting that decides how eager merging is.

In the course's LlamaIndex version, a hierarchical parser produces the levels (for example 2048, 512 and 128 tokens). Only the smallest are embedded for matching, and the larger ones are kept in a store so they can be returned.

4Side by side

When each fits.

Sentence-windowAuto-merging
Matches onOne sentenceA small chunk
Expands contextAlways, by a fixed windowOnly when enough siblings match
Relevant lines far apartNot connectedConnected through the shared parent
Set-upLight: one parser and a window sizeHeavier: levels, a store and a threshold
Good whenThe answer sits in a tight neighbourhoodEvidence is scattered through a section

The course presents them as complements, not rivals. Both are tuned to a document type, since a contract and an invoice would want different structures.

5Measure before you trust it

Compare scores, not impressions.

"It reads better" is not evidence. The course scores each variant with three checks: is the retrieved context relevant to the question, is the answer grounded in that context, and does it answer the question. Change one setting, rerun the same questions, and compare scores. Evaluating RAG shows how.

Not always worth it. If your documents are short and uniform, plain chunks plus a reranker may be enough. Add machinery only when you can point to a failure it fixes.

6What has changed since the course

Current versus old.

In the course materialsStatus todayUse now
LlamaIndex ServiceContextRemovedUse the global Settings object
Imports from the top-level llama_index packageMovedImport from llama_index.core and its sub-packages
Retired chat models used in the demosRetiredPick a current model from your provider
The ideas: sentence-window and auto-mergingStill soundSame mechanism, new code
Net read: the retrieval ideas aged well. Only configuration objects, import paths and model names moved. Check the current LlamaIndex docs before copying any course code.

7Check yourself

Short questions and one line to remember.

What does sentence-window retrieval do?

Matches on single sentences, then returns each hit with its neighbouring sentences.

What does auto-merging retrieval do?

Matches on small chunks and replaces them with their parent when enough of that parent's children are retrieved.

Why not just use big chunks?

A big chunk mixes topics, so its embedding matches less precisely. These methods match small and expand afterwards.

Try it: edit the threshold in the auto-merging snippet to 0.8 and rerun. Which sections still merge, and why?
One line to remember: "Match on small pieces, then give the model the bigger piece around them."
Notes based on DeepLearning.AI’s Building and Evaluating Advanced RAG with Jerry Liu and Anupam Datta. Library APIs change, so check the current LlamaIndex documentation. Independent notes, not affiliated with or endorsed by any company or course named here. Examples and wording are my own.