What is GraphRAG?
Retrieval that uses structure rather than similarity alone. Two quite different approaches share the name, and conflating them is the most common source of surprise in a budget review, because one is nearly free to set up and the other is not.
Separating them before anything else makes the rest of the subject tractable.
| Approach | What it does | What it needs |
|---|---|---|
| Traversal | Retrieves by walking relationships in a graph you already maintain, then passes what it gathered to the model | An existing graph. Cheap to query, and the graph has to come from somewhere |
| Constructed | Builds the graph from your documents with a language model, partitions it into communities, summarises each one, then answers from those summaries | A language model pass over the entire corpus. This is where the cost lives |
The second is what the original research described and what most tooling implements. The distinction matters commercially because a vendor demonstrating traversal over a graph they already built for you is showing something quite different from one proposing to construct a graph from your corpus, and only the second carries an indexing bill. The structure itself is covered at knowledge graph.
What can GraphRAG do that vector RAG cannot?
Answer questions about the corpus as a whole. The original paper is precise on this: the contrast with vector retrieval is the ability to handle queries requiring global sensemaking across the entire dataset. That is a narrower claim than the multi-hop framing used to sell it, and a more defensible one.
The reason is structural rather than a matter of quality. Retrieval returns the top matching chunks, so a question like what are the recurring themes across these ten thousand contracts has no correct set of chunks to return. The answer is a property of everything, and retrieval by definition hands back a subset.
Constructed GraphRAG solves that by precomputing the whole-corpus view: it partitions the graph into thematic communities and summarises each, so a global question is answered from summaries rather than from retrieved passages.
- Global questions. Themes, patterns, what changed across a body of documents, what a whole corpus implies. This is the genuine capability gap and nothing in vector retrieval addresses it.
- Multi-hop questions. Real, and less exclusive than claimed. Vector retrieval can often reach a multi-hop answer given enough chunks and a capable model, so the advantage here is one of reliability rather than possibility.
- Point lookups. No advantage at all. Where one passage contains the answer, benchmark work consistently finds vector retrieval matches or beats graph approaches at a fraction of the cost.
Worth separating from the general graph question. Whether to hold a graph at all is decided by the traversal test at knowledge graph: does the answer exist in one place, or only in a chain. GraphRAG is a narrower decision about retrieval technique, and it only arises once you have already decided the structure is worth having. The wider retrieval loop, including why retrieval rather than generation is the current bottleneck, is at agentic RAG.
How does the indexing actually work?
Four stages, and the first is the expensive one. A language model reads every chunk and extracts entities and relationships. Those become a graph. A clustering algorithm partitions the graph into communities. A model then summarises each community, and those summaries become the index that global queries read.
Understanding the stages matters because the two query modes read different parts of the output.
- Local search starts from entities mentioned in the question and traverses outward, gathering neighbouring entities, relationships and the source chunks. This is what answers a specific question about particular things.
- Global search ignores the question's entities and reads community summaries at a chosen level of the hierarchy. This is what answers a question about the corpus, and it is the mode that costs the most to run.
The community layer is the part people skip when explaining this, and it is the whole mechanism. Communities are a thematic partitioning of the graph produced by clustering, not a taxonomy anybody designed, which means their quality depends entirely on whether the clustering found coherent themes. Published critiques note that detected communities often mix themes, which degrades the summaries and therefore the global answers.
Lighter variants exist and are worth knowing about. Some simplify the extraction to a dual entity and relation level; others defer the summarisation entirely, building the index as cheaply as vector retrieval and doing more work at query time. The published figures for that deferred approach put indexing cost at roughly a thousandth of the full method with comparable global answer quality.
Where does the cost sit?
At indexing, and it inverts the usual retrieval economics. Vector retrieval is cheap to index and cheap to query. Full constructed GraphRAG is expensive to index, expensive to query globally, and expensive again every time the corpus changes enough to warrant a rebuild.
Two measurements from the literature give the shape of it better than any percentage.
- Global querying is not a rounding error. One peer-reviewed analysis reported that on a standard multi-hop benchmark the method detected nearly three thousand communities, and that answering one hundred questions consumed on the order of a hundred million tokens. The authors described that overhead as impractical.
- Indexing is a function of corpus size, not query volume. Which means it is knowable in advance, and should be a specific number in any proposal rather than a percentage estimate.
- The rebuild cadence decides everything. If documents change daily or weekly, indexing is a recurring subscription rather than a one-time investment. This is the single question most likely to be missed in a business case.
- Entity drift adds a hidden cost. The same person or company arrives as several entities across documents, and resolving that is ongoing work rather than a setup task.
Does GraphRAG beat vector RAG?
On the right question shape, clearly. In general, no, and the honest evidence here is stronger than the marketing. A 2026 benchmark study found that graph-based retrieval frequently underperforms conventional retrieval on many real-world tasks, with performance highly dependent on the domain.
That finding deserves to sit next to the capability claim rather than below it, because both are true and they apply to different questions.
| The claim | What the evidence actually shows |
|---|---|
| Better on global questions | Supported, and it is the original paper's own framing. Nothing in vector retrieval addresses whole-corpus sensemaking |
| Better on multi-hop | Supported on reliability, with reported recall improvements. Vector retrieval frequently reaches the same answer given enough context |
| Better generally | Not supported. A 2026 benchmark reports frequent underperformance against conventional retrieval on real-world tasks |
| Better on lookups | Contradicted. Graphs add cost and noise where one passage holds the answer |
The practical consequence is a procurement rule rather than an architectural one. A benchmark on news articles or research papers does not predict behaviour on your contracts or your claims files, because the technique's advantage depends on whether your documents actually contain the dense entity relationships the extraction step is looking for. Demand a proof of concept on your own corpus, and treat any vendor benchmark as evidence the method works somewhere rather than evidence it will work here.
How should you decide whether to use GraphRAG?
Decide in two steps. First, look at the kind of question people ask. If answers need connections across many facts, GraphRAG is relevant. Second, look at how often your data changes. Frequent changes make the map costly to keep current. Most teams check cost first and are surprised later.
Five checks, in the order that eliminates fastest.
-
Classify a real sample of your questions
Take fifty questions people actually ask and sort them into lookups, multi-hop and global. If nothing lands in global and few land in multi-hop, the decision is made and vector retrieval is the answer.
-
Establish how often the corpus changes materially
Daily or weekly change makes full construction a recurring cost. Quarterly or slower makes it an investment. This single number moves the economics more than any other input.
-
Ask for the indexing cost as a figure, not a ratio
It is a function of corpus size and therefore knowable in advance. A proposal that expresses it as a percentage of something else is deferring a number somebody will discover later.
-
Consider the lighter variants before the full method
If the global capability is what you want but the indexing cost is prohibitive, the deferred approaches reach comparable global quality at a small fraction of the setup cost by moving work to query time.
-
Run the proof of concept on your documents
Not a sample corpus and not the vendor's benchmark. The technique's advantage is domain-dependent in a way the published results make unusually clear.
One architectural note that applies whatever you choose. The strongest production systems route rather than pick, sending lookups to vector retrieval and reserving graph traversal for the questions that need it, because the cost difference between the two paths is large enough to make routing worth the complexity. That routing decision is the same shape as the one at model routing, and it carries the same requirement: record which path served each question, or you cannot tell whether the graph earned its place.