GraphRAG vs Vector RAG: The 1,000x Index Bill Is the Small One

GraphRAG's index costs about 1,000x a vector index, but at 2,000 questions a day its global search path costs far more than the index. Start with hybrid search and add a graph only where your eval set shows global questions failing.

By Rajesh Beri·October 9, 2026·14 min read
Share:
A long conference table covered in thousands of printed index cards connected by red string, with a single slim card-catalogue drawer of neatly filed cards at the near end of the table.

Illustration generated using AI

If someone has told you a knowledge graph will fix your retrieval, the indexing invoice is the number they quoted and the query invoice is the one that will hurt. Microsoft Research's own figures put vector RAG indexing at 0.1% of the cost of full GraphRAG, a 1,000x gap. An independent benchmark then measured GraphRAG's global search using an average of 331,375 tokens per question against 879 for plain vector RAG, 377 times as many, on every query.

The verdict: run hybrid search (dense vectors plus BM25 keyword search plus a reranker) as your default, and add a graph only when your own evaluation set shows a class of global or multi-hop questions that the hybrid stack fails and that someone will pay to answer. If you do add one, start with LightRAG or a deferred-summarisation design, not Microsoft GraphRAG's standard index. Full GraphRAG is the loser here for anything that changes weekly: its open-source repository now says it is "largely in maintenance mode", and its best feature is also its most expensive one.

The workload behind every number below: a 10M-token corpus, 10% of it changing each month, 2,000 questions a day, with gpt-6.1-sol at $2 input / $10 output per 1M tokens and text-embedding-3-large at $0.13 per 1M (prices checked on OpenAI's live pricing page on October 10, 2026).

Option Index build, 10M tokens (our model) Prompt tokens per question (GraphRAG-Bench) Answers whole-corpus questions? Who should pick it
Hybrid vector + BM25 + reranker ~$1.40 879 (measured on plain vector RAG) No Everyone, first
LightRAG ~$330 100,832 Partly Teams whose eval proves a multi-hop gap, on Postgres
Microsoft GraphRAG, standard ~$430 38,707 local / 331,375 global Yes Static corpora with low query volume
Microsoft GraphRAG, fast ~$100 Same query paths as standard Yes, noisier A cheap first test of global questions
LazyGraphRAG Same as vector RAG (Microsoft's claim) Microsoft claims 700x below global search Yes Azure shops that can use Microsoft Discovery
Neo4j as the graph store $0.09 per GB-hour, Professional tier Depends on your pipeline Depends on your extraction Teams with a real ontology to model

Which Questions Can Vector RAG Not Answer?

Vector RAG cannot answer a question whose answer is spread across the whole corpus rather than sitting in a few passages. Microsoft's original GraphRAG post named the two failure classes in February 2024: baseline RAG "struggles to connect the dots" across separate documents, and it "performs poorly when being asked to holistically understand summarized semantic concepts" over a large collection. "What are the main themes in 5,000 customer complaints?" is the canonical example. Top-k retrieval returns the 10 most similar chunks, and no 10 chunks contain a summary of 5,000 documents.

That class is real. It is also much smaller than the pitch implies. Most enterprise questions are lookups: what does clause 14.2 say, which SKU did this customer buy, what is the escalation path for a P1. GraphRAG-Bench, a 2025 benchmark revised in February 2026, found that "Basic RAG Matches GraphRAG in simple fact retrieval". On its novel-text dataset, reranked vector RAG scored 60.92 on fact retrieval against 49.29 for Microsoft GraphRAG. A separate unified comparison by Zhou et al. found the same shape on PopQA: vanilla RAG at 60.8 accuracy, GraphRAG's local search at 45.5.

The graph wins where the question needs synthesis. On GraphRAG-Bench's complex-reasoning tasks over novels, reranked RAG scored 42.93 while Microsoft GraphRAG scored 50.93. On contextual summarisation, 51.30 against 64.40. Then on the medical dataset's summarisation task, reranked RAG scored 65.75, far ahead of Microsoft GraphRAG at 41.87 but behind Fast-GraphRAG at 67.88. The advantage depends on the task and on the corpus, which is why the decision has to come from your own question mix.


What Does GraphRAG Indexing Actually Cost?

GraphRAG indexing costs roughly 300x to 2,000x a vector index at our workload, because an LLM reads every chunk at least twice and then writes a report for every community it finds. A vector index only embeds each chunk once.

Microsoft's standard method "uses a language model for all reasoning tasks": entity extraction, relationship extraction, description summarisation and community reports. The shipped defaults chunk text at 1,200 tokens with 100 overlap, run one extra "gleaning" pass per chunk, and allow each community report 8,000 tokens of input and 2,000 of output.

Our model of the 10M-token build on gpt-6.1-sol, with assumptions stated so you can swap in your own:

  • Extraction covers about 9,100 chunks, each sent twice (the initial pass plus the default gleaning), at an assumed 8,000 input and 2,000 output tokens per chunk in total. That is 73M input tokens and 18M output, about $328.
  • Community reports: LightRAG's authors counted 1,399 communities on a 5.1M-token legal corpus. Scaled linearly, that is about 2,750 here, each at up to 8,000 in and 2,000 out, about $99.
  • The total is about $430, before description summaries, which add more.

The vector index embeds about 10.9M tokens once: $1.42 with text-embedding-3-large, $0.22 with text-embedding-3-small at $0.02 per 1M. Against the small embedding model the ratio is close to 2,000x, against the large one about 300x, and Microsoft's 1,000x sits between them.

Swap in gpt-6-luna at $0.10 / $0.50 per 1M and the graph build drops to about $21. That is the most effective single lever, and Microsoft's own README warns that "GraphRAG indexing can be an expensive operation" and to "start small." The catch is that extraction quality falls with model quality, and the graph is only as good as the extraction (more on that below).

At $430, the index is not the problem for most enterprises. The two bills that are, re-indexing and querying, never show up in the demo.


Why Re-Indexing Is the Recurring Line Item

Re-indexing cost tracks how much of your corpus changes, not how big it is, because entity extraction re-runs on every changed document and community structure can shift with each update. Model a year of churn before you approve the first build.

GraphRAG 1.0 added an update command that "computes the deltas between an existing index and newly added content" and relies on an LLM cache so re-runs are "often significantly faster and cheaper than an initial run." The CLI reference lists standard-update and fast-update methods. Microsoft's documentation describes this for newly added content. It does not spell out how edited or deleted documents are handled, so test exactly that before you trust it with a policy library that is revised in place.

The community layer is where the cost hides. The LightRAG paper's analysis is that GraphRAG must regenerate its community structure on an incremental update, at around 1,399 x 2 x 5,000 tokens for their legal corpus. LightRAG instead inserts new entities and relationships into the existing graph without rebuilding communities.

At 10% monthly churn on our workload:

  • Vector RAG re-embeds 1M tokens, about $0.14 a month. A year costs about $3.10 including the first build.
  • GraphRAG standard re-extracts 1M tokens (about $33), plus between nothing and all ~$99 of community reports depending on how far the changes reach. That is $33 to $132 a month, and $830 to $2,010 for year one.

None of these numbers would worry a CFO. The problem is that they scale with churn: a support knowledge base rewritten weekly, or a contract repository with daily redlines, multiplies them, and a stale graph silently answers from last month's facts.


What Does Each Query Cost, and How Long Does It Take?

Query cost is the bill that dominates, because graph methods put far more retrieved text into each prompt. GraphRAG-Bench measured average prompt tokens per question on its novel dataset: 879 for vector RAG, 1,008 for HippoRAG2, 38,707 for GraphRAG local search, 100,832 for LightRAG, and 331,375 for GraphRAG global search.

At 2,000 questions a day (60,000 a month), priced as input tokens on gpt-6.1-sol:

Method Prompt tokens per question Input cost per question Input cost per month
Vector RAG 879 $0.0018 ~$105
GraphRAG local 38,707 $0.077 ~$4,600
LightRAG 100,832 $0.20 ~$12,100
GraphRAG global 331,375 $0.66 ~$39,800

Those token counts come from the benchmark's configuration, not yours, and output tokens are left out. The ratios are what transfer. Treat the global figure as the total across global search's map-reduce calls, not one prompt: GraphRAG-Bench's own text says a single global-search prompt reaches up to 40,000 tokens, while its table averages 331,375 per question. Zhou et al. measured the same order independently, about 300,000 tokens per global query on their MultihopSum set. Microsoft's own query documentation calls global search, which runs map-reduce over every community report, "a resource-intensive method." LazyGraphRAG exists largely because of it: Microsoft says LazyGraphRAG reaches comparable quality to global search on global queries at "more than 700 times lower query cost".

One discrepancy is worth knowing before a LightRAG vendor deck reaches you. The LightRAG paper reports "fewer than 100 tokens for keyword generation and retrieval", while GraphRAG-Bench measured 100,832 prompt tokens per question. Both can be true: the first counts the retrieval call, the second counts the context sent to the answering model. Your invoice is the second.

Latency follows the tokens. Zhou et al. measured 2.35 seconds per multi-hop question for vanilla RAG and 19.28 seconds for LightRAG. On abstract questions, vanilla RAG took 18.7 seconds and GraphRAG global search 72.2 seconds, and on their MultihopSum set "each query in GGraphRAG takes around 9 minutes." An analyst tool can absorb 72 seconds. A support agent's copilot cannot.


How Do the Cheaper Graph Variants Compare?

Each cheaper variant moves LLM work out of indexing, and each gives up something specific in exchange.

GraphRAG fast replaces LLM extraction with "noun phrases extracted using NLP libraries such as NLTK and spaCy," builds relationships from co-occurrence within text units, and skips entity descriptions. The LLM still writes community reports, so our model puts the build near $100. Microsoft says plainly that "the graph tends to be quite a bit noisier" and that fast mode suits global summary questions, while standard suits detailed entity exploration. Do not pick it if you plan to query the graph directly or reuse it outside GraphRAG.

LazyGraphRAG defers all LLM summarisation to query time and spends a tunable "relevance test budget" (Microsoft tested 100, 500 and 1,500) on the slice of the graph each question needs. Its indexing cost is, per Microsoft, "identical to vector RAG", and at a 500 budget its query cost was 4% of GraphRAG's community-level-2 global search. The evidence has limits: one dataset of 5,590 AP news articles, 100 synthetic queries, and an LLM judge. A June 2025 editor's note on the same post says LazyGraphRAG went into Microsoft Discovery and Azure Local; the open-source repository does not mention it. Do not plan around it unless you can buy it through those products.

LightRAG keeps a knowledge graph and a vector index side by side and skips community reports entirely. It is MIT-licensed, has about 40,000 GitHub stars, and offers five query modes (local, global, hybrid, naive and mix, the default). Its own README says the default file-based storage is "not for production" and recommends PostgreSQL, which also covers vectors and graph. Extraction cost is close to GraphRAG's, about $330 in our model. Do not pick it if your documents are edited in place often: the README documents incremental insertion and deletion with graph regeneration, so plan for each edited document to cost a delete plus a re-insert, and measure that before production.

HippoRAG2, a research system, is worth one line because GraphRAG-Bench measured it at 1,008 prompt tokens per question, close to vector RAG, while scoring 53.38 on complex reasoning over novels. It is the evidence that graph retrieval does not have to carry GraphRAG's query bill.


Graph Quality Depends on the Ontology You Build

The quality of a GraphRAG answer is decided by extraction accuracy, entity resolution and deduplication, and none of those comes with the database licence. A graph where "IBM", "International Business Machines" and "Big Blue" are three nodes answers a question about IBM's contracts with a third of the evidence.

Neo4j's own GraphRAG course is candid about it. Its builder by default merges only entities with "the same label and identical name property". Fuzzy matching (RapidFuzz) and semantic matching (spaCy) resolvers exist, and the course warns that post-processing carries "the risk of incorrectly merging distinct entities." Without resolution, you get "multiple nodes representing the same real-world entity." Merge two subsidiaries with similar names into one node and the graph will attribute one company's contracts to the other.

The store itself is cheap. Neo4j AuraDB Professional is $0.09 per GB per hour with a 1GB minimum, about $66 a month per GB, and Business Critical is $0.20 per GB per hour with a 2GB minimum, about $292 a month (prices checked October 10, 2026). Memgraph and Postgres are alternatives LightRAG supports. The expensive part is the person-weeks to define 10 to 20 entity types around the questions your users ask, write the resolution rules, and review a sample of merges. If nobody owns the ontology, do not buy the graph.


What the Hybrid Baseline Already Buys You

A hybrid stack of dense vectors, BM25 keyword search and a reranker closes most of the gap that teams try to fix with a graph. Anthropic's contextual retrieval study cut the top-20 retrieval failure rate from 5.7% to 1.9% by combining contextual embeddings, BM25 and reranking, a 67% reduction. GraphRAG-Bench's best vector configuration was the one with the reranker.

Any of Qdrant, Weaviate, Pinecone or Postgres with pgvector and full-text search can serve it, and our vector database comparison covers that choice. The strongest case for the graph is that hybrid search will never summarise 5,000 documents. That is true, and a scheduled job that pre-computes summaries for the five aggregate questions your executives actually ask costs a fraction of a community index and is easier to audit.


The Decision: What Predicts Regret

Nobody has published a win rate on your documents. Every number above comes from a corpus someone else chose: AP news for Microsoft, novels and medical texts for GraphRAG-Bench, a legal set for LightRAG. Build the question set first.

The criteria that predict regret:

  1. The share of global questions. If fewer than one in ten real questions needs whole-corpus synthesis, a graph is a cost center; pre-compute answers to those few.
  2. Document churn. Above roughly 10% a month, model the update bill and test edits and deletions as well as additions.
  3. Query volume times prompt size. Multiply your daily questions by the per-question token counts above. That product sets your budget.
  4. Latency tolerance. A copilot has seconds and an analyst has minutes, and global search runs in minutes.
  5. An ontology owner. If no named person owns entity types and merge rules, skip the graph.

This Week: pull 200 real questions from search logs or tickets and label each as lookup, multi-hop or global. Our guide to building a RAG eval set by hand covers the method.

This Month: run the hybrid baseline against that set and record failures by label. Only the multi-hop and global failures are candidates for a graph.

Before You Sign the Indexing Job: index a 5% sample with GraphRAG fast and LightRAG, measure prompt tokens per question on your failed questions, and multiply by a year of query volume and churn.

The Bottom Line

The graph-retrieval pitch in 2026 resembles the data-lake pitch of 2014: a real capability for a narrow class of questions, sold as the default architecture. Teams that approve the $430 index without modeling the $40,000-a-month global search path will discover the second number in their cloud bill.

Price the query path on your own question mix before you approve the index.

Continue Reading

Share:

Frequently Asked Questions

How much more does GraphRAG cost to index than vector RAG?

Microsoft Research puts vector RAG indexing at 0.1% of full GraphRAG's cost, about 1,000x. On a 10M-token corpus with gpt-6.1-sol, a modeled GraphRAG build costs about $430 against $1.42 to embed the same text with text-embedding-3-large (prices checked October 10, 2026).

When does GraphRAG beat vector RAG?

On global and multi-hop questions that need synthesis across many documents, such as summarising themes in thousands of records. On simple fact lookups, GraphRAG-Bench found reranked vector RAG scored 60.92 against 49.29 for Microsoft GraphRAG.

What does GraphRAG cost per query?

GraphRAG-Bench measured 38,707 prompt tokens per question for GraphRAG local search and 331,375 for global search, against 879 for vector RAG. At 2,000 questions a day on gpt-6.1-sol, global search input alone comes to about $39,800 a month.

Is LightRAG cheaper than Microsoft GraphRAG?

LightRAG skips community reports, so its index build and incremental updates are cheaper, but its entity extraction costs about the same. Per question, GraphRAG-Bench measured 100,832 prompt tokens for LightRAG, more than GraphRAG local search.

Can I use LazyGraphRAG in the open-source GraphRAG package?

Microsoft's June 2025 note says LazyGraphRAG was integrated into Microsoft Discovery and Azure Local. The open-source microsoft/graphrag repository does not mention it and describes itself as largely in maintenance mode.

Newsletter

Stay Ahead of the Curve

Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.

Subscribe

Latest Articles

View All →