Indexing
Once a document is chunked, each chunk becomes a vector and gets stored somewhere built to search by meaning instead of by exact match.
Chunks to embeddings
one chunk of text
Each chunk that comes out of splitting gets converted into an embedding: a fixed-length vector of numbers, produced by a model trained so that chunks with similar meaning end up with similar vectors, regardless of which exact words they use.
This is a completely separate model from the LLM that eventually generates an answer. Its only job is to turn text into a point in a high-dimensional space, hundreds or thousands of dimensions, where distance between points stands in for difference in meaning.
The vector store
indexed for fast nearest-vector lookup, not just exact match
Every chunk's embedding gets stored in a vector store, indexed in a way that makes it fast to find the vectors nearest to any new point, not just fast to look one up by exact match.
A question gets embedded with the exact same model used for the chunks, landing it in that same space. That's what makes the next step, retrieval, work at all: without a shared embedding space, comparing a question to a document would be meaningless.
Most vector stores also keep the original chunk text and any metadata, like source document or page number, alongside the vector, so a nearby vector can be turned back into an actual passage of text once it's found.