Skip to main content
Concept
GraphRAG is a form of retrieval-augmented generation (RAG) that retrieves context from a graph of entities and relationships, usually alongside the source text, instead of only from the text chunks most similar to the question. The graph lets the retriever follow connections between facts, which helps with questions that span several documents, center on one entity, or ask about a whole collection. In exchange, it adds an indexing step that extracts entities and relationships, typically with model calls, and needs ongoing work to keep the graph accurate.

Learning objectives

After reading this article you will be able to:
  • Explain how GraphRAG retrieves context through entities and relationships, not only similar chunks
  • Recognize the question types where vector-only RAG falls short
  • Compare GraphRAG with vector RAG on retrieval unit, indexing cost, freshness, and explainability
  • Decide when GraphRAG is worth its extraction and maintenance cost

Where does vector RAG fall short?

Vector RAG falls short when the answer depends on how facts in different passages connect, because similarity search ranks each chunk on its own. Standard vector RAG splits documents into chunks, turns each chunk into a vector embedding, and puts the chunks most similar to the question into the model’s prompt, which works well when the answer sits in one or two passages that are semantically close to the question. For the full pipeline, see What is retrieval-augmented generation (RAG)? Chunking also strips context. A passage that says “it was rolled back” may not say what “it” refers to.

How does GraphRAG work?

GraphRAG builds a graph from the same content and uses it at query time. Common patterns include:
  • Entity and relationship extraction. At indexing time, a model or NLP pipeline reads each chunk and emits entities, such as people, products, services, and contracts, plus typed relationships between them. Each chunk links to the entities it mentions.
  • Entity resolution. Mentions that refer to the same thing, such as “the billing service” and “billing-svc”, are merged into one node so the facts about it connect.
  • Graph expansion from retrieved nodes. Vector or keyword search finds entry points, and a bounded traversal collects neighboring entities, relationships, and the chunks that support them. This is sometimes called local search.
  • Graph-scoped retrieval. A traversal defines the candidate set first, such as the documents linked to one customer or product area, and similarity search runs inside it. This is a form of filtered vector search.
  • Community or cluster summaries. A community detection algorithm partitions the entity graph into groups of densely connected entities, and a model writes a summary of each group ahead of time. Questions about the whole collection are answered by combining the relevant summaries. This is sometimes called global search.

How does GraphRAG compare with vector RAG?

Vector RAG retrieves chunks by similarity; GraphRAG also retrieves entities and relationships by traversal, at higher indexing and maintenance cost. GraphRAG follows edges for multi-hop questions, gathers the facts linked to a resolved entity, answers global questions from cluster summaries when they are built, and can model permissions as edges that retrieval traverses.

What does GraphRAG cost?

GraphRAG costs more than vector RAG to build and maintain: extraction at indexing time, work to keep the graph accurate and fresh, and bounds on query fan-out.
  • Extraction cost. With model-based extraction, every chunk goes through a model at indexing time, and changing the extraction prompt or model can mean reprocessing the whole collection.
  • Graph quality. Extraction can miss relationships or invent ones that are not in the text. Entity resolution fails in both directions: duplicates split one entity’s facts across several nodes, and over-merging combines different entities into one.
  • Freshness. Entities and summaries are shared across documents, so an edit to one document can change nodes, edges, and summaries that other documents also rely on.
  • Consistency. When the graph, vector index, and keyword index live in separate systems, the application must keep them in sync.
  • Query fan-out. Highly connected entities can pull in large neighborhoods, so expansion needs bounds; the FAQ below covers hop limits.
Constraining entity and relationship types with a schema, keeping an edge from every extracted fact back to its source chunk, and evaluating against a fixed question set keep these costs visible.

When should you use GraphRAG?

GraphRAG pays off when questions are about connections: multi-hop questions, entity profiles, dependency and impact analysis, or themes across a collection. It also fits when the domain already has structure, such as customers, products, tickets, and contracts, or when access rules are relationships that retrieval must respect. Vector or hybrid RAG is often enough when answers live in single passages, as with FAQs and product documentation, or when content changes faster than extraction can keep up. A practical path is to start with hybrid vector and BM25 retrieval, collect the questions it answers poorly, and add graph structure where those failures cluster.

How do you implement GraphRAG?

You implement GraphRAG by adding entity extraction and graph-aware retrieval on top of a standard chunk-and-embed pipeline, then checking it against vector-only retrieval.
1

Define the schema

Choose entity types, relationship types, and how documents and chunks are stored.
2

Index the content

Chunk documents, compute embeddings, and index chunk text for vector and keyword search.
3

Extract and resolve

Extract entities and relationships from each chunk, merge duplicate entities, and link each chunk to the entities it mentions.
4

Retrieve with the graph

Find entry points, traverse a bounded neighborhood, search within it, and fuse the results into context with citations.
5

Evaluate

Compare answers with vector-only retrieval on multi-hop, entity, and summary questions.

How does HelixDB support GraphRAG?

HelixDB is a graph database with native vector search and BM25 full-text search, so documents, chunks, entities, and relationships can live in one labeled property graph. It is a multigraph, so two entities can be connected by several labeled edges. Vector and text indexes on chunk properties sit beside the graph rather than in a separate system. Retrieval can traverse first and then search within the traversal. With prefiltering, the order is graph traversal, exact candidate membership, ranking, then top k, and a result outside the candidate set is never returned. A prefiltered search accepts up to 1,000,000 unique candidates and returns at most 800 results. Traversals support bounded repeats and shortest paths for expansion. Each request is one ACID transaction over a committed snapshot, so the traversals, vector searches, and text searches in a request read the same data. The application computes embeddings, and HelixDB stores and indexes the vectors. HelixDB has no built-in rank fusion or reranking, so the application combines the vector and BM25 results.

Frequently asked questions

How many hops should graph expansion follow?

Many systems start with one or two hops from each entry point. Each hop can multiply the number of nodes reached, and highly connected entities can fill the context with loosely related facts. Capping the neighbors followed per node and filtering by relationship type keep expansion focused; raise the limits only when evaluation shows missed answers.

Do I need a knowledge graph before using GraphRAG?

You need a graph, but it does not have to be built by hand. Many pipelines extract it from documents with a model, and many domains already have structured relationships, such as ticket links, product catalogs, or ownership records, that can seed it. See What is a knowledge graph? They are complementary. Hybrid search combines vector and keyword ranking to find good entry points, and GraphRAG adds traversal from those entry points. GraphRAG systems can use hybrid search as their first retrieval step. An agent memory store can be modeled as a graph of facts, entities, sources, and versions scoped to one user. Recalling from it uses the same moves as GraphRAG: find relevant facts by similarity and keywords, then expand through relationships. See What is AI agent memory?

What is retrieval-augmented generation (RAG)?

The standard chunk, embed, retrieve, and generate pipeline.

What is a knowledge graph?

Entities, relationships, and provenance for search and LLMs.

What is hybrid search?

Combine vector similarity with BM25 keyword ranking.

What is filtered vector search?

Rank only the candidates that pass a filter or traversal.

How do you build long-term memory for AI agents?

Apply graph-plus-search retrieval to per-user agent memory.

Do you need separate graph, vector, and text databases?

Trade-offs of keeping all three access paths in one system.

Prefiltered search

Rank only the nodes a traversal reaches in a HelixDB query.

Traversals

Follow outgoing and incoming relationships in a HelixDB query.