1Purpose first
Escape the chunk-size choice.
Plain retrieval forces one choice of chunk size. Small chunks match well but give the model little to work with. Big chunks give context but match poorly. The techniques here refuse the choice: search on small pieces, then hand the model something bigger.
2Sentence-window retrieval
Hit plus neighbours.
Index single sentences so matching is precise. When a sentence matches, return it together with the sentences around it.
doc = ["Refunds are allowed within 30 days.", "Items must be unused.",
"Shipping costs are not refunded.", "Digital goods cannot be refunded.",
"Contact support to start a refund."]
def window(sents, i, w=1):
# return the matched sentence plus w sentences either side
return " ".join(sents[max(0, i - w): i + w + 1])
hit = 2 # say the search matched this sentence
print("matched :", doc[hit])
print("returned :", window(doc, hit))
print("wider :", window(doc, hit, w=2))matched : Shipping costs are not refunded. returned : Items must be unused. Shipping costs are not refunded. Digital goods cannot be refunded. wider : Refunds are allowed within 30 days. Items must be unused. Shipping costs are not refunded. Digital goods cannot be refunded. Contact support to start a refund.
The window is fixed, so it only helps when the answer sits in the neighbouring lines. If the relevant sentences are far apart, it will not connect them.
3Auto-merging retrieval
Merge children into the parent.
Build a hierarchy: small chunks (children) inside larger sections (parents). Index and match the small ones. If enough children of the same parent are retrieved, swap them for the parent. The more of a section that matches, the more of it you get.
# small chunks (children) sit inside larger sections (parents)
parents = {
"refunds": ["Refunds are allowed within 30 days.", "Items must be unused.", "Refunds go to the original payment method."],
"shipping": ["Shipping takes 3 to 5 days.", "Tracking is sent by email.", "Remote areas can take longer."],
}
child_of = {c: p for p, kids in parents.items() for c in kids}
def auto_merge(retrieved, threshold=0.5):
out, by_parent = [], {}
for c in retrieved:
by_parent.setdefault(child_of[c], []).append(c)
for p, hits in by_parent.items():
if len(hits) / len(parents[p]) >= threshold:
out.append("[merged] " + " ".join(parents[p])) # enough siblings matched: return the parent
else:
out.extend(hits) # otherwise keep the small hits
return out
retrieved = ["Refunds are allowed within 30 days.", "Items must be unused.", "Tracking is sent by email."]
for line in auto_merge(retrieved): print(line)[merged] Refunds are allowed within 30 days. Items must be unused. Refunds go to the original payment method. Tracking is sent by email.
Two of the three refund sentences matched, which is at least half, so the whole refund section was returned. Only one shipping sentence matched, so it stayed small. The threshold is the setting that decides how eager merging is.
In the course's LlamaIndex version, a hierarchical parser produces the levels (for example 2048, 512 and 128 tokens). Only the smallest are embedded for matching, and the larger ones are kept in a store so they can be returned.
4Side by side
When each fits.
| Sentence-window | Auto-merging | |
|---|---|---|
| Matches on | One sentence | A small chunk |
| Expands context | Always, by a fixed window | Only when enough siblings match |
| Relevant lines far apart | Not connected | Connected through the shared parent |
| Set-up | Light: one parser and a window size | Heavier: levels, a store and a threshold |
| Good when | The answer sits in a tight neighbourhood | Evidence is scattered through a section |
The course presents them as complements, not rivals. Both are tuned to a document type, since a contract and an invoice would want different structures.
5Measure before you trust it
Compare scores, not impressions.
"It reads better" is not evidence. The course scores each variant with three checks: is the retrieved context relevant to the question, is the answer grounded in that context, and does it answer the question. Change one setting, rerun the same questions, and compare scores. Evaluating RAG shows how.
6What has changed since the course
Current versus old.
| In the course materials | Status today | Use now |
|---|---|---|
LlamaIndex ServiceContext | Removed | Use the global Settings object |
Imports from the top-level llama_index package | Moved | Import from llama_index.core and its sub-packages |
| Retired chat models used in the demos | Retired | Pick a current model from your provider |
| The ideas: sentence-window and auto-merging | Still sound | Same mechanism, new code |
7Check yourself
Short questions and one line to remember.
What does sentence-window retrieval do?
Matches on single sentences, then returns each hit with its neighbouring sentences.
What does auto-merging retrieval do?
Matches on small chunks and replaces them with their parent when enough of that parent's children are retrieved.
Why not just use big chunks?
A big chunk mixes topics, so its embedding matches less precisely. These methods match small and expand afterwards.