Chunking

The first real decision in a RAG pipeline, made before anything gets embedded or indexed: how a document gets cut into the pieces everything downstream will actually operate on.

Compare real chunking strategies on your own text
01

Why split, and how big

whole document, embedded as one vector

one vector, the average of everything in it

Embedding models have a limited context window, commonly 512 to 8,000 tokens, so a whole document rarely fits into one embedding call. Splitting is also what makes retrieval precise: embed an entire 30-page document as one vector and it captures the average of everything in it, not the one paragraph that actually answers a question.

Chunk size is a real tradeoff. Too small and a chunk loses the surrounding context that gave it meaning, a fragment about "the policy" with no antecedent for what policy. Too large and the chunk dilutes relevance, compressed into one vector alongside a lot of text that has nothing to do with the question, and costs more tokens once it's stuffed into a prompt.

Overlap fixes the boundary problem specifically: instead of every chunk starting exactly where the last one ended, each new chunk repeats a slice of the previous chunk's tail. A sentence that would otherwise get cut in half by an unlucky boundary still appears whole in at least one chunk.

02

Splitting strategies

fixed character count, blind to content

The simplest splitter just cuts every N characters or tokens: fast, predictable, and blind to what it's cutting through. A table row or a sentence gets severed exactly as easily as the whitespace between two paragraphs.

Recursive character splitting is the practical default. It tries to cut on paragraph breaks first, falls back to sentence breaks, then words, only reaching for a hard character cut as a last resort, keeping most chunks intact along natural boundaries without needing to understand the content.

Semantic splitting goes further: it embeds individual sentences and cuts wherever the similarity between consecutive sentences drops, wherever the topic actually shifts, instead of wherever a fixed count runs out. It produces more meaningful chunks at the cost of an embedding pass just to decide where to split.