Chunking strategy comparison

Paste your own text, pick a target size and an overlap, and see how three real splitting strategies cut it differently — including how often each one slices a sentence in half, and whether overlap actually stitches it back together in the next chunk (checked, not assumed).

Settled on a setting? chunk_size and chunk_overlap below map directly onto LangChain's real splitters — copy the config or the chunks themselves.

180 chars
90 chars

Each chunk after the first repeats this many characters from the end of the one before it — shown highlighted below. Turned up here so the effect is visible; real pipelines typically use less (10–20% of chunk size).

Fixed-size

Cuts every N characters. Fast, predictable, blind to content.

7

Chunks

6

Mid-sentence cuts

5/6

Recovered by overlap

180–270 chars, avg 252

CharacterTextSplitter(chunk_size=180, chunk_overlap=90, separator="")
1 · 180 chars
Retrieval-augmented generation grounds a language model's answer in documents it was never trained on. A user's question is embedded into the same vector space as the document chun
2 · 270 chars · 90 repeated
trained on. A user's question is embedded into the same vector space as the document chunks, and the chunks closest to that question vector are retrieved and stuffed into the prompt alongside it. A chunk boundary that happens to fall in the middle of the one sentence
3 · 270 chars · 90 repeated
pt alongside it. A chunk boundary that happens to fall in the middle of the one sentence that actually answers the question breaks retrieval entirely. A retriever that returns a chunk containing only half of that sentence gives the model half an answer to reason from,
4 · 270 chars · 90 repeated
hunk containing only half of that sentence gives the model half an answer to reason from, and the model has no way to know the other half exists somewhere else. Different splitters make this failure more or less likely. A splitter that cuts every N characters, blind to
5 · 270 chars · 90 repeated
s make this failure more or less likely. A splitter that cuts every N characters, blind to what it's cutting through, will sever a sentence exactly as often as it severs whitespace between two paragraphs. A splitter that respects sentence boundaries never produces that
6 · 270 chars · 90 repeated
between two paragraphs. A splitter that respects sentence boundaries never produces that specific failure, at the cost of chunks whose length varies with however long the sentences happen to be. A splitter that tries paragraph breaks first, then sentences, then words,
7 · 231 chars · 90 repeated
s happen to be. A splitter that tries paragraph breaks first, then sentences, then words, only reaching for a hard character cut as a last resort, sits between the two: mostly intact boundaries, but still a predictable target size.

Recursive

Tries paragraph breaks, then sentences, then words, before a hard cut.

11

Chunks

3

Mid-sentence cuts

0/3

Recovered by overlap

75–270 chars, avg 175

Every cut here is a single sentence longer than the whole chunk — overlap can't repair that (it would need to repeat almost the entire previous chunk). A bigger chunk size would fix this; more overlap won't.

RecursiveCharacterTextSplitter(chunk_size=180, chunk_overlap=90)
1 · 102 chars
Retrieval-augmented generation grounds a language model's answer in documents it was never trained on.
2 · 270 chars · 90 repeated
gmented generation grounds a language model's answer in documents it was never trained on. A user's question is embedded into the same vector space as the document chunks, and the chunks closest to that question vector are retrieved and stuffed into the prompt alongside
3 · 94 chars · 90 repeated
chunks closest to that question vector are retrieved and stuffed into the prompt alongside it.
4 · 138 chars · 3 repeated
it. A chunk boundary that happens to fall in the middle of the one sentence that actually answers the question breaks retrieval entirely.
5 · 263 chars · 90 repeated
e middle of the one sentence that actually answers the question breaks retrieval entirely. A retriever that returns a chunk containing only half of that sentence gives the model half an answer to reason from, and the model has no way to know the other half exists
6 · 106 chars · 90 repeated
odel half an answer to reason from, and the model has no way to know the other half exists somewhere else.
7 · 75 chars · 15 repeated
somewhere else. Different splitters make this failure more or less likely.
8 · 222 chars · 58 repeated
Different splitters make this failure more or less likely. A splitter that cuts every N characters, blind to what it's cutting through, will sever a sentence exactly as often as it severs whitespace between two paragraphs.
9 · 261 chars · 90 repeated
gh, will sever a sentence exactly as often as it severs whitespace between two paragraphs. A splitter that respects sentence boundaries never produces that specific failure, at the cost of chunks whose length varies with however long the sentences happen to be.
10 · 269 chars · 90 repeated
e, at the cost of chunks whose length varies with however long the sentences happen to be. A splitter that tries paragraph breaks first, then sentences, then words, only reaching for a hard character cut as a last resort, sits between the two: mostly intact boundaries,
11 · 127 chars · 90 repeated
for a hard character cut as a last resort, sits between the two: mostly intact boundaries, but still a predictable target size.

Sentence-packed

Packs whole sentences up to the target size. Never cuts one in half.

8

Chunks

0

Mid-sentence cuts

0/0

Recovered by overlap

102–306 chars, avg 227

No mid-sentence cuts here, so there's nothing for overlap to repair — the repeated text below is still applied, just with nothing broken to fix.

1 · 102 chars
Retrieval-augmented generation grounds a language model's answer in documents it was never trained on.
2 · 274 chars · 90 repeated
gmented generation grounds a language model's answer in documents it was never trained on. A user's question is embedded into the same vector space as the document chunks, and the chunks closest to that question vector are retrieved and stuffed into the prompt alongside it.
3 · 225 chars · 90 repeated
ks closest to that question vector are retrieved and stuffed into the prompt alongside it. A chunk boundary that happens to fall in the middle of the one sentence that actually answers the question breaks retrieval entirely.
4 · 279 chars · 90 repeated
e middle of the one sentence that actually answers the question breaks retrieval entirely. A retriever that returns a chunk containing only half of that sentence gives the model half an answer to reason from, and the model has no way to know the other half exists somewhere else.
5 · 150 chars · 90 repeated
wer to reason from, and the model has no way to know the other half exists somewhere else. Different splitters make this failure more or less likely.
6 · 222 chars · 58 repeated
Different splitters make this failure more or less likely. A splitter that cuts every N characters, blind to what it's cutting through, will sever a sentence exactly as often as it severs whitespace between two paragraphs.
7 · 261 chars · 90 repeated
gh, will sever a sentence exactly as often as it severs whitespace between two paragraphs. A splitter that respects sentence boundaries never produces that specific failure, at the cost of chunks whose length varies with however long the sentences happen to be.
8 · 306 chars · 90 repeated
e, at the cost of chunks whose length varies with however long the sentences happen to be. A splitter that tries paragraph breaks first, then sentences, then words, only reaching for a hard character cut as a last resort, sits between the two: mostly intact boundaries, but still a predictable target size.