GraphRAG & Hybrid Search

Keyword search finds exact words. Semantic search finds similar meaning. Neither finds a relationship no one ever wrote down, that is what adding a knowledge graph buys a retrieval pipeline.

01

Why retrieval exists: the context window problem

A 30-page PDF100k tokens

Every LLM has a context window, a hard limit on how much text fits into a single call. A 30-page PDF runs to roughly 100,000 tokens; feed that straight into a model with a 50,000-token window and half of it simply never arrives.

Even a generous window doesn't solve this generally. Llama 3.1 shipped with 128k tokens of context, large enough for that one PDF, but most models are smaller, and reprocessing an entire document on every single question doesn't scale even when the window technically fits.

So instead of pasting in everything, a system retrieves only the slice that's actually relevant to the question being asked. How a document gets split into pieces and turned into something searchable, chunking strategy, embedding model choice, which vector database to use, is its own deep topic for later; what matters here is the layer above that: once you have retrievable pieces, which strategy actually finds the right one.

02

Keyword search: exact, sparse, and brittle

sentenceTheLionRoars
"The lion roars"111

The oldest retrieval strategy is keyword search: bag-of-words or TF-IDF, techniques that reduce a sentence to which words it contains. "The lion roars" becomes a row of 1s and 0s over a fixed vocabulary, one column per word, a 1 wherever that word appears.

Do this across a whole vocabulary and most of every row is zero, only the words actually used get marked. That's called a sparse matrix, and training anything on data this sparse tends to overfit: the model learns specific words instead of the meaning behind them.

Keyword search is fast, exact, and genuinely necessary for the cases it handles well, an invoice number, a product code, a name spelled a specific way. What it cannot do is match meaning: search "vacation days" against a document that only says "PTO" and keyword search finds nothing.

03

Semantic search: similarity in vector space

KingQueenManWoman

Semantic search, also called dense vector search, takes the opposite approach. Every chunk of text gets embedded, converted into a vector of numbers that captures meaning rather than exact wording, using models trained so that similar meanings end up close together in that space.

The classic illustration: embed "King," "Queen," "Man," and "Woman," and King lands nearer to Queen than to either Man or Woman, the same relationship that holds in meaning holds in the vector space, without a single letter overlapping between the words.

A query gets embedded the same way, then compared against every stored vector using cosine similarity, and the closest matches come back regardless of exact wording. This is what finds "vacation days" against a document that says "PTO": the two phrases sit close together in vector space even though they share no words.

04

Hybrid search: best of both

Keyword searchcatches exact IDs, codes, names
Vector searchcatches synonyms, paraphrases

Keyword search catches exact terms and misses synonyms. Semantic search catches synonyms and can blur past an exact code or ID that needed to match precisely. Neither one alone is reliable enough on its own for a production system.

Hybrid search runs both and merges the results, so an exact invoice number and a paraphrased policy question both come back correctly from the same query. Most production vector databases, Pinecone, Astra DB, FAISS, Chroma, support this combination directly rather than requiring two separate pipelines bolted together.

This is usually where a RAG system's accuracy actually comes from, not a bigger model, not a longer prompt, but giving retrieval two independent ways to find the right passage instead of one.

05

Add the graph: GraphRAG

Keyword searchsparse, exact terms (TF-IDF)
Vector searchdense, semantic similarity
Graph traversalstructured, follows relationships

Hybrid search still only answers what's textually or semantically similar. It has no way to answer what's connected to this, because neither keyword nor vector search understands relationships, only content. That's exactly the gap a knowledge graph fills as a third retrieval path.

LangChain's Neo4jGraph connects to a running graph database, and LLMGraphTransformer can build that graph automatically, fed raw text, letting an LLM extract the nodes and relationships instead of someone writing Cypher by hand for every document. GraphCypherQAChain does the reverse: given a natural-language question, an LLM writes the Cypher query to answer it, runs it against the graph, and returns the structured result.

Put all three together and a single question gets answered by whichever path actually holds the answer: an exact term found by keyword search, a paraphrased concept found by vector search, or a relationship that only exists as a path through the graph. That combination, keyword, semantic, and graph, folded into one retrieved context, is what turns plain RAG into GraphRAG.

GraphRAG isn't one single technique either. Generating Cypher at query time, the pattern above, works well for precise questions with a clear entity to anchor on, "who are Rohit Sharma's teammates." A second pattern builds the graph once, clusters it into densely connected communities, and has an LLM summarize each community ahead of time, which is what actually answers a broad, thematic question like "what are the top five themes in this dataset," since no single Cypher query could return a theme that only emerges from reading across hundreds of documents at once.

Try it: which retrieval method finds it?

A query, and what the source text actually says. Guess whether keyword search, semantic search, or a graph traversal is the one that actually retrieves the answer.

"Find the document mentioning invoice INV-88231."

The right document exists, and it contains that exact string.

Which retrieval method actually finds this?

Waiting for your guess.