Concept
AI agent memory is the infrastructure that lets an agent built on a large language model
(LLM) keep information beyond a single model call: what a user told it, what happened in
earlier sessions, and how it should carry out recurring tasks. Because the model itself
is stateless, memory lives in an external store that the agent writes to as it learns
something and reads from when it needs context, typically before generating a reply.
Learning objectives
After reading this article you will be able to:
- Explain why a model’s context window cannot serve as an agent’s long-term memory
- Describe the main types of agent memory, from working memory to the user profile
- Outline the memory lifecycle from ingestion through forgetting and recall
- Identify the storage and retrieval requirements a memory layer must meet
Why is a context window not memory?
A context window is not memory because it holds only what a model processes in one call, and it is cleared when that call ends. The window is a maximum amount of text, measured in tokens; in many models the reply the model generates counts against the same limit. It works like working memory: everything the model knows about the current task has to be in it. Treating it as long-term memory breaks down for several reasons:- It resets. A new session starts empty unless something outside the model puts earlier information back.
- It is bounded and typically priced per token. Replaying every past conversation grows cost and latency with each turn, and eventually stops fitting.
- It has no notion of truth over time. If a user said one thing last month and the opposite today, a transcript contains both, and the model has to guess which is current.
- It is not scoped. Nothing in a raw prompt records who owns a fact, where it came from, or whether it may be shown to this user.
- More text is not always better context. Models do not use every part of a long prompt equally well, so a focused set of relevant facts often produces better answers than a full history.
What types of memory do AI agents use?
AI agents use short-term memory (working memory and conversation history) and long-term memory (episodic, semantic, and procedural), often condensed into a user profile. These terms are borrowed from cognitive psychology. Short-term memory covers what the agent is working with right now; long-term memory persists across sessions.
The difference between episodic and semantic memory matters in practice. Episodes
answer “what happened”; semantic facts answer “what is true now.” Agents often derive
semantic facts from episodes, for example inferring a preference from several
conversations, and keep a link back to those episodes as evidence. The user profile
is not a separate source of truth. It is a summary maintained from long-term memories so
the agent can load broad context cheaply on every request.
How does the memory lifecycle work?
The memory lifecycle is a loop of six steps (ingest, extract, deduplicate, update, forget, and recall): each interaction can add, change, or remove memories, and each request reads from the result.- Ingest. Collect raw input: conversation turns, uploaded documents, tool results, and application events.
- Extract. Turn raw input into atomic, self-contained facts. A reply such as “next Tuesday” only means something together with the question before it, so the extractor needs recent conversation context and the current date.
- Deduplicate. Compare each candidate with existing memories, by meaning and by exact terms, so the same fact is not stored many times.
- Update and version. When a fact changes, write a new version and mark the old one as superseded instead of overwriting it. History stays available for audits and for questions about what changed.
- Forget. Expire time-bound facts, let low-value episodes decay, and delete what a user or policy asks to remove.
- Recall. Retrieve the memories this user may see, that are still current, and that are relevant to the request, then assemble them into context.
What infrastructure does a memory layer need?
A memory layer needs tenant-scoped identity, relationships between records, semantic recall, exact recall, lifecycle filtering, and provenance. Each lifecycle step places a different demand on storage and retrieval:
Many teams assemble these from several systems, such as a relational store for records,
a vector store for embeddings, and a search engine for keywords. That works, but the
application then has to keep them consistent. A memory deleted in one system can still
surface from another, and a permission filter applied after a similarity search can
remove every result or, if it is forgotten, leak one. See
one database for graph, vector, and text
for the trade-offs.
How does HelixDB support agent memory?
HelixDB is an open-source graph database with native vector search and BM25 full-text search, so the requirements above can be met in one labeled property graph:- Properties with indexes hold tenant scope, identity, and lifecycle fields such as
isLatest,validTo,deletedAt, andexpiresAt. Equality and range secondary indexes make them indexed filters. - Edges record provenance, categories, entities, version updates, and associations.
- Vector search handles deduplication and paraphrase recall, and BM25 handles exact names, IDs, and rare tokens. Both index types can be partitioned by tenant.
- A profile node holds stable user context that the agent loads on every request.
Frequently asked questions
How does an agent decide which memories to recall?
It narrows by scope first: only memories owned by or shared with the current user or team. It then keeps only current memories, excluding anything deleted, superseded, or expired, and ranks what remains by relevance with semantic and keyword search. These filters should apply before or during ranking, not after, and the agent keeps only a small budget of top results so the context stays focused.Is agent memory the same as RAG?
They share retrieval techniques but differ in what they retrieve. Retrieval-augmented generation (RAG) typically reads from a document collection that changes independently of the agent. Memory is written by the agent from its own interactions, changes as facts are corrected, and is scoped to a user or team. Many systems combine the two and link each memory to the source passage it came from.Can a vector database alone serve as agent memory?
Not on its own. A vector database covers semantic recall, and many also filter on metadata copied onto each vector, which can handle ownership and lifecycle flags. That copy has to be kept in sync, and it cannot represent provenance, versions, or links between entities as relationships; many vector databases also lack keyword ranking for exact identifiers. A complete memory layer needs properties, lifecycle filters, keyword search, and relationships alongside the vectors.What should an agent store in memory?
Store information that is likely to matter again: stable facts and preferences, notable events, and workflows the agent has learned. Skip small talk, transient details, and anything a user or policy says must not be kept. Each stored fact should be atomic and self-contained, and should record its owner, its source, and when it was learned so it can be superseded or expired later.How should an AI agent forget?
Many systems use soft deletion: mark a memory as deleted, superseded, or expired and exclude it from every recall path, and reserve physical deletion for user requests and retention policies. The guide to building long-term memory covers expiry, decay, and deletion in detail.Related topics
Building long-term memory
Data model, write path, forgetting, and recall for agent memory.
What is RAG?
Ground a model’s answers in data retrieved at query time.
What is GraphRAG?
Retrieval over entities and relationships, not only text chunks.
What is hybrid search?
Combine vector similarity with BM25 keyword ranking.
What is filtered vector search?
Restrict similarity search to the records a user may see.
What is a vector database?
Store embeddings and find the nearest ones to a query.