Concept
A knowledge graph is a structured collection of facts about entities, such as people,
products, documents, and concepts, and the typed relationships between them, typically
organized by a schema or ontology. Facts link entities to each other or to attribute
values, and each fact can record where it came from, so software can answer questions by
following connections and cite its sources. Knowledge graphs are commonly stored in
graph databases and used to ground search, recommendations, and large language model
(LLM) applications in connected, verifiable facts.
Learning objectives
After reading this article you will be able to:
- Describe the parts of a knowledge graph, from entities to provenance
- Compare a knowledge graph with a graph database
- Explain how a knowledge graph is built, including extraction and entity resolution
- Describe how a knowledge graph is kept current as sources change
What is in a knowledge graph?
A knowledge graph contains entities, the typed relationships between them, attributes that describe both, and provenance for each fact, usually organized by an ontology or schema that defines the allowed types.
A schema or ontology is usually what separates a knowledge graph from an arbitrary set
of links. It fixes what kinds of things exist and how they may relate, so facts from
different sources land in the same shape and queries mean the same thing everywhere. An
ontology usually goes further than a schema: it can also define class hierarchies,
constraints, and inference rules.
What is the difference between a knowledge graph and a graph database?
A knowledge graph is data with agreed meaning. A graph database is software that stores and queries graph-shaped data. The two are often used together but are not the same.
A knowledge graph can be stored as a
property graph or as RDF triples. See
property graph vs RDF for how the two
models differ.
How do you build a knowledge graph?
You build a knowledge graph by defining an ontology, extracting entities and relationships from your sources, resolving duplicate entities, recording provenance, and loading the result into a database.1
Define the ontology
List the entity types, relationship types, and required attributes your questions
need. Start small; a focused schema is easier to extract into and to query.
2
Collect sources
Structured records such as account and product tables map to entities almost
directly. Unstructured sources such as documents, tickets, and conversations need
extraction.
3
Extract entities and relationships
For unstructured text, the application runs an extraction pipeline: split each
document into chunks, ask an LLM or another extraction model to return the entities
and relationships it finds in a fixed format that matches the ontology, and reject
output that does not validate against the schema.
4
Resolve entities
Decide when two mentions refer to the same thing. One ticket says “the billing
page” and another says “the invoices screen”; both may mean one feature. Typical
techniques combine exact identifiers, normalized names, embedding similarity to find
candidates, and rules or review to confirm a merge.
5
Record provenance
Link each extracted fact to the chunk and document it came from, with the extraction
time and a confidence value. Provenance supports citations, audits, and later
corrections.
6
Load, index, and query
Write entities, relationships, and provenance together, and add indexes for the
lookups and searches your application runs.
How do you keep a knowledge graph fresh?
You keep a knowledge graph fresh with a maintenance loop rather than a one-time load: ingest changes incrementally, supersede outdated facts, retract facts by source, and re-run entity resolution.- Ingest incrementally. Track which facts came from which source, and re-extract only the documents that changed.
- Supersede rather than overwrite. Mark outdated facts with validity fields such as
validToorisLatest, and link a new version to the one it replaces, so history stays queryable. - Retract by source. When a document is deleted or corrected, provenance tells you exactly which facts to remove or re-check.
- Re-run entity resolution. New mentions can reveal that two existing entities are the same, or that one entity should be split.
How do knowledge graphs help LLMs and RAG?
A knowledge graph lets RAG retrieve connected, typed facts rather than only passages similar to the question, so it can answer multi-hop questions and cite sources. Retrieval-augmented generation (RAG) supplies an LLM with retrieved context before it answers, and retrieval based only on vector similarity struggles with questions such as “which customers reported problems with features in the Reports product?”, where the answer is spread across entities and relationships rather than contained in one passage. Following explicit facts across the graph is an approach often called GraphRAG. LLM-extracted facts can be wrong, so keep confidence values and provenance, and let downstream prompts treat low-confidence facts accordingly.How does HelixDB store a knowledge graph?
HelixDB stores data as one labeled property graph, which maps directly onto a knowledge graph:- Entities and relationships. Entities are nodes and relationships are directed edges, each with exactly one label and typed properties. Because HelixDB is a multigraph, two sources asserting the same relationship can be kept as separate edges with their own properties.
- Embeddings and text on the same nodes. Store an embedding and a text field on an entity or chunk node, then add a vector index and a text index for similarity and BM25 search. The application computes the embeddings; HelixDB stores and indexes them.
- Provenance edges. Connect entities, or facts modeled as their own nodes, to their source chunks and documents with edges that carry properties such as confidence or extraction time. For a fact stored as a relationship edge, record provenance as properties on that edge, such as the source chunk ID, confidence, and extraction time. Indexes can target edge properties as well as node properties.
- Identity and freshness. Unique equality indexes on node labels suit canonical
identifiers, and equality or range indexes on fields such as
isLatestorvalidTomake queries that filter to current facts efficient. - Scoped retrieval. Vector and BM25 search can be prefiltered to a traversal-defined candidate set, such as documents a user can read.
Frequently asked questions
What is the difference between a knowledge graph and an ontology?
An ontology is a formal specification of a domain: the types of entities and relationships and the rules that govern them, such as class hierarchies and constraints. The knowledge graph’s instance data, its entities and facts, follows the ontology, and many systems store the ontology in the graph alongside that data. One ontology can describe many knowledge graphs.Do you need RDF to build a knowledge graph?
No. RDF is one standard way to represent a knowledge graph, and it suits publishing and integrating data across organizations. A labeled property graph works as well and is common for application-facing knowledge graphs. See property graph vs RDF.Can an LLM build a knowledge graph automatically?
An LLM can extract candidate entities and relationships from text at scale, but the output needs validation against the ontology, entity resolution, and provenance. Treat extraction as one stage of a pipeline the application controls.How is a knowledge graph different from a vector database?
A vector database retrieves items by similarity of embeddings. A knowledge graph represents explicit, typed facts and how they connect. Many AI applications use both: vectors to find relevant starting points, and the graph to follow facts from there. See What is a vector database?.Related topics
What is a graph database?
Nodes, edges, traversals, and when a graph is the right model.
What is a property graph?
Labeled nodes and edges with typed properties on both.
Property graph vs RDF
Two ways to model a knowledge graph, compared side by side.
What is RAG?
Grounding a model’s answer in data retrieved at query time.
What is GraphRAG?
Retrieval that follows relationships, compared with vector RAG.
What is AI agent memory?
Working, episodic, and semantic memory for LLM agents.