6 Best Zep Alternatives for Temporal and Multi-Agent Memory in 2026
Where Zep stops fitting — the end of self-hosting, per-byte write costs, and silent fact invalidation — and the six alternatives that each win a narrower job.
October 5, 2026 · Caura.AI
The best Zep alternative for teams running several agents, or for anyone who wants time-aware memory without operating a graph database, is Caura. It records when each fact took effect, answers “what was true in March” with As-Of Recall, supersedes stale facts, and runs on PostgreSQL under Apache 2.0.
If you want Zep’s exact temporal graph and can run Neo4j or FalkorDB yourself, Graphiti, the open-source engine under Zep, is the only complete like-for-like replacement. Mem0, Hindsight, Cognee, and Supermemory each fit a narrower reason for leaving, covered below.
Zep alternatives covered:
- Caura: governed shared memory with event-time validity, As-Of Recall and supersession, on Postgres.
- Graphiti (self-hosted): Zep’s own temporal knowledge graph engine, minus the managed service.
- Mem0: the simplest drop-in memory API, with timestamps on every fact.
- Hindsight: a single-container memory engine with a temporal retrieval strategy and no graph database.
- Cognee: a knowledge graph built from documents, code, and business data.
- Supermemory: a memory API with plugins for coding agents.
Why do teams choose Zep in the first place?
Zep solved a problem most memory layers ignore: facts change. In Zep, every fact is an edge in a knowledge graph with a validity window, so the graph knows that a customer used the Starter plan until June and the Pro plan after it. The design is written up in the paper Zep: A Temporal Knowledge Graph Architecture for Agent Memory, and the engine is open source as Graphiti, which has 31.2k GitHub stars as of 30 September 2026.
Three other things keep teams on Zep. Retrieval, storage, and users are unmetered on Zep Cloud, so a read-heavy agent pays only for what it writes. Custom entity and edge types let a healthcare or finance team model its own domain. And Zep ships Python, TypeScript and Go SDKs, which few memory vendors match.
Those strengths are real. People search for Zep alternatives because of what sits around them.
Why are Zep users looking for alternatives?
Six problems show up in Zep’s own documentation, pricing, and the Graphiti issue tracker. Most of them get worse as message volume or agent count grows.
1. Self-hosting ended with Community Edition
Zep deprecated its self-hosted Community Edition in April 2025, and the getzep/zep repository now marks it unsupported. Teams that need on-prem or air-gapped memory have two paths: Zep’s Enterprise BYOC deployment, or running Graphiti directly. Graphiti needs a graph database (Neo4j 5.26, FalkorDB or Amazon Neptune with OpenSearch), an LLM provider and an embedder.
Graphiti’s Kuzu backend is marked deprecated in the Graphiti README, after the Kuzu project itself was archived in October 2025. That is three systems to provision and monitor where Community Edition used to be one Docker container.
2. Every write fans out into serial LLM calls
Graphiti turns each episode into a graph by calling an LLM several times: once to extract entities, again to resolve each entity, again for each fact, and again to deduplicate and invalidate against the existing graph.
The write path is where Zep’s accuracy comes from, and where its cost and latency live.
One user measured 626 seconds for a one-sentence episode on a rate-limited endpoint and projected 30 to 50 minutes for a 5 KB document. Another reported spending about $0.80 for 40 short chats on default OpenAI models, roughly two cents per conversation before any retrieval. The README notes that Graphiti works best with models that support structured output, and local-model setups hit validation errors, as in issue #868 where the Ollama quickstart fails on a missing pydantic field.
3. Fact invalidation can retire facts that are still true
Invalidation is Zep’s headline feature, and it is also where the sharpest recent bug report sits. In issue #1728, a team found that 1,616 of 3,950 facts in its production graph (41%) carried an invalid_at date. A hand audit of four of them found three were collateral: saving a memory that merely mentioned an entity retired an unrelated fact about it, because edge invalidation searched the whole graph. The retired facts stay in the graph but drop out of search, so the loss is silent.
Issue #1707 describes a second silent failure, where add_memory returns success and the episode is later discarded.
These reports come from self-hosted Graphiti. Zep Cloud may run a different pipeline, so test with your own data before assuming the same rate.
4. The managed price moved, and compliance is Enterprise-only
As of 30 September 2026, Zep Cloud bills by bytes written. One credit covers an episode up to 350 bytes, plus one more credit per additional 350 bytes. The Flex plan is currently $125 a month for 50,000 credits, with overage at $25 per 10,000. Flex Plus is $375 for 200,000 credits. Zep’s plan comparison reserves SOC 2 Type II, a HIPAA BAA, and audit logs for Enterprise.
At the 700-byte message size Zep’s own estimator uses, Flex Plus becomes the cheaper plan above 75,000 messages a month.
A support agent handling 300 conversations a day at 12 messages each writes about 108,000 messages a month. Under Zep’s listed 40,000-credit and 10,000-credit top-up blocks, that is 216,000 credits, or $450 a month on Flex Plus and $550 on Flex. Several 2026 comparison posts still quote Flex at $25, so budget from the live pricing page.
5. The API keeps changing under you
Zep has changed its API or its deployment model at four points so far.
Each change was defensible on its own. Together they add migration work to every year of a Zep deployment.
Zep 1.0 moved memory onto Graphiti and deprecated custom facts and the Collections endpoints. The February 2026 deprecation wave deprecated the V2 SDK, renamed sessions to threads and groups to graphs, removed fact ratings, and removed session.extract and session.classify with no replacement.
6. Benchmark numbers that did not survive review
Zep’s original 84% LoCoMo result was challenged by Mem0’s CTO in getzep/zep-papers issue #5, who re-ran it at 58.44%. Zep corrected its own figure to 75.14% and argued Mem0 had misconfigured the test. In the same correction, Zep noted that a plain full-context baseline scored about 73%. Neither side is neutral, but the episode is a reason to ask any vendor, Zep included, for runnable evaluation code.
Where is Zep actually strong, and what fully replaces it there?
Zep is the strongest option in the category for one job: answering questions about how a subject’s facts changed over time. Any honest shortlist starts by deciding whether you need that job done, and by whom.
Bi-temporal facts on a single subject
Zep tracks both when a fact was true and when it was recorded. Graphiti is the only complete replacement, because it is the same engine. You keep validity windows, custom ontologies and graph traversal, and you take on the database, the LLM bill on writes and the invalidation behavior described above.
Time-aware recall without a graph database
If the question you need answered is “what was true on this date” or “what is true now”, Caura covers it on Postgres. Each memory carries an event time (ts_valid_start) separate from its ingest time; recall accepts a valid_at date, and supersession retires the old value when a new one arrives. As-Of Recall then measures freshness from the question’s date, so an imported two-year history keeps its real timeline.
Unmetered reads for read-heavy agents
Zep’s billing favors agents that read far more than they write. Self-hosted options (Caura, Graphiti, Hindsight, Cognee) remove per-read cost entirely. Among managed plans, Caura Pro includes 50,000 searches a month for $49.
How do the best Zep alternatives compare?
The table scores each of the six Zep alternatives on the reasons teams leave.
| Tool | How it handles changed facts | Self-host footprint | Write-path cost | Multi-agent governance | License |
|---|---|---|---|---|---|
| Caura | Event-time validity, As-Of Recall, supersession | Postgres, pgvector, Redis | One enrichment pass per write | Trust tiers, keystones, audit log | Apache 2.0 |
| Graphiti | Bi-temporal edges, invalidation | Graph DB plus LLM and embedder | Several LLM calls per episode | None built in | Apache 2.0 |
| Mem0 | Keeps old and new facts with timestamps | Vector store | One extraction pass | Agent ID scoping | Apache 2.0 |
| Hindsight | Temporal retrieval strategy | One container, embedded Postgres | Extraction on retain | Per-bank isolation | MIT |
| Cognee | Graph updated as data changes | Embedded stores by default | Graph build per ingest | Deployment permissions | Apache 2.0 |
| Supermemory | Temporal changes and forgetting | Self-host or cloud | Extraction per add | Per-container isolation | MIT |
1. Caura: temporal, governed memory without a graph database
Repository: github.com/caura-ai/caura
Caura is the Zep alternative for teams that want time-aware memory shared across agents, on infrastructure they already know how to run. An agent writes plain text with caura_write. A single LLM pass classifies it, extracts entities into a knowledge graph, scans for PII, checks for contradictions, and stamps its scope. Any authorized agent recalls it with hybrid search across vectors, keywords, and graph hops.
How Caura answers each reason teams leave Zep
Self-hosting: The open-source release is the full engine: storage, contradiction detection, supersession, audit trail, and 12 MCP tools on PostgreSQL with pgvector and Redis, started with docker compose up. A local embedder profile runs with zero outbound API calls for air-gapped sites.
Write-path cost: Enrichment is one LLM pass per write, and it is included on every plan, including Free. There is no per-byte meter.
Changed facts: Supersession marks the older memory as outdated instead of deleting it, and every transition is audit-logged, so a wrong retirement is visible and traceable. The caura-long-run-fleet reference shows it on a 14-day simulation: a competitor price holds at $299 for eight days, changes to $349 on day 9, and eight stale memories are superseded so the day 10 brief recalls one number. The walkthrough is in The price changed. Your agents didn’t notice.
Governance across agents: Zep scopes memory to users and graphs. Caura adds per-agent trust levels: standard agents read and write in their own fleet, cross-fleet agents can read wider, and only admin-level agents can delete memories. Keystones serve mandatory policy rules to every agent at session start. The caura-cross-fleet-gov demo enforces fleet boundaries as a SQL predicate on every recall.
Pricing: Every plan allows unlimited agents, fleets, and users. As of 30 September 2026, Pro is $49 a month ($41 billed annually) for 250,000 memories, 25,000 writes, and 50,000 searches, listed on the pricing page.
If your Zep bill is driven by message volume and you run more than one agent, Caura’s Zep comparison maps each capability, and the free tier connects to Claude Code, Cursor or any MCP client with one config block.
Caura results on the questions Zep is known for
Knowledge-update and temporal-reasoning questions test exactly the behavior Zep users pay for.
Caura scores 92.2% on LongMemEval overall (461 of 500), with 97.4% on knowledge-update questions and 91.0% on temporal reasoning, using one configuration and a 22.4k-token median context. In a separate warm, single-tenant benchmark, search ran at 23 ms p50. The evaluation code and per-question verdicts are public in caura-longmemeval. In production, eToro’s Company Brain runs 300+ agents on Caura with 26,500+ memories.
Where Caura is weaker: it does not offer Zep’s custom entity and edge ontology, and its graph is extracted automatically rather than modeled. The community is small next to Graphiti’s.
Best for: teams leaving Zep over cost or self-hosting who run more than one agent, or who need an audit trail on every memory change.
2. Graphiti: Zep’s temporal graph, self-hosted
Repository: github.com/getzep/graphiti
Graphiti is the Zep alternative that changes the least, because it is Zep’s own engine. You keep bi-temporal edges, custom entity types, and hybrid search across semantic, keyword, and graph traversal. You lose Zep Cloud’s context blocks, managed scaling, and compliance attestations, and you inherit every issue in sections 2 and 3 above: serial LLM calls on writes, a hard dependency on structured-output models, and invalidation you need to monitor.
Best for: teams with graph database experience who need validity windows exactly as Zep models them and want to stop paying per byte.
3. Mem0: the simplest swap away from Zep’s complexity
Repository: github.com/mem0ai/mem0
Mem0 replaces Zep’s episodes, threads, and context templates with two calls, add() and search(). Its v3 algorithm keeps both the old and new versions of a changed fact with timestamps, which gives the model temporal context but leaves choosing the current value to your code, as users describe in issue #4956. Entity linking sits on the $249 Pro plan, and the self-hosted release dropped external graph stores.
Best for: single-agent personalization where Zep’s modeling depth was more than you needed.
4. Hindsight: self-hosted memory with no graph database
Repository: github.com/vectorize-io/hindsight
Hindsight runs as one Docker container with embedded Postgres, which is the shortest path from Zep Community Edition to something you run yourself. Each query runs several retrieval strategies, including a temporal one, and merges them with a reranker. Its README states that its LongMemEval results were independently reproduced by Virginia Tech’s Sanghani Center. Memory is organized into banks with retain, recall, and reflect operations. Banks isolate agents but do not govern what they share.
Best for: teams that lost Community Edition and want a single self-hosted service with strong recall.
5. Cognee: graphs from documents and business data
Repository: github.com/topoteretes/cognee
Some teams chose Zep for its JSON and business-data ingestion more than for conversation memory. Cognee fits that reason better. It builds a self-hosted knowledge graph from documents, code, tickets, and conversations, defaults to embedded stores for local use, and ships an MCP server. It models relationships across a corpus rather than validity windows on a subject, so test time-sensitive questions before migrating.
Best for: agents that reason over existing documentation and records.
6. Supermemory: memory for coding agents and assistants
Repository: github.com/supermemoryai/supermemory
Supermemory bundles fact extraction, user profiles, and document retrieval behind one API, with plugins and an MCP server for Claude Code, Cursor, Codex, and OpenCode. Its README says it handles temporal changes, contradictions, and automatic forgetting, and reports first place on LongMemEval, LoCoMo, and ConvoMem. Those results are vendor-reported, so run your own questions first.
Best for: developers who want persistent memory inside a coding agent with the least setup.
How do you choose the right Zep alternative?
Start from the reason you are leaving Zep, then answer four questions before you migrate a single episode.
Do you need “what was true then”, or only “what is true now”?
Write a fact, replace it, then ask for both the current value and the value on an earlier date. Graphiti and Caura answer both inside the memory layer. Mem0 returns both versions and leaves the choice to you. Most agents only need the current value, and paying for a full temporal graph to get it is the most common overspend in this category.
Who will operate the graph database?
If nobody on your team runs Neo4j or FalkorDB today, Graphiti adds that job. Caura, Hindsight, and Cognee run without one.
How many agents write to the same memory?
Zep’s model is one graph per subject. Once two or more agents write about the same customer, you also need to know which agent wrote a fact, what each agent may read, and what happens when they disagree. Caura’s guide to agent fleet memory covers the scope, provenance, trust, and validity fields that make shared writes safe.
What does ingestion cost at your real message volume?
Multiply daily messages by average bytes, divide by 350 for Zep credits, and compare against self-hosted LLM costs for the same writes. Caura’s deterministic memory write-up also shows how Zep, Mem0, and Letta each decide when a write happens, which drives the bill as much as the price per write.
Which Zep alternative should you pick?
Pick Caura if more than one agent shares memory, if you need self-hosting on Postgres, or if you want time-aware recall without a per-byte bill. It is the only option here with event-time validity, supersession, and per-agent governance in one engine, and it runs in production at 300+ agents.
Pick Graphiti if you need Zep’s exact bi-temporal graph and have the team to operate it. For a single agent, pick Mem0 for the simplest API, Hindsight for a one-container self-host, Cognee when your knowledge lives in documents, and Supermemory for coding agents.
If you are on Zep today, the fastest test is to point two agents at Caura’s free tier through its MCP server, write a price from one agent, change it, and ask the other agent for both today’s price and last week’s. The answers show whether you still need a graph database to get Zep’s best feature.
Frequently Asked Questions
Can I still self-host Zep?
Not as Zep. Community Edition was deprecated in April 2025 and is unsupported. You can self-host Graphiti, Zep’s open-source engine, with your own graph database, or use Zep’s Enterprise BYOC option inside your VPC.
Is Graphiti the same as Zep?
Graphiti is the temporal knowledge graph engine that powers Zep. Zep Cloud adds managed hosting, context blocks, user and thread management, and compliance attestations on top. Self-hosting Graphiti means building those parts yourself.
How much does Zep cost in 2026?
As of 30 September 2026, Zep Cloud is free for 10,000 credits a month, then $125 a month on Flex for 50,000 credits or $375 on Flex Plus for 200,000. One credit covers an episode up to 350 bytes. Retrieval and storage are not metered, and SOC 2 Type II, HIPAA BAA and audit logs are Enterprise features.
Which Zep alternative handles facts that change over time?
Graphiti keeps Zep’s bi-temporal model exactly. Caura records event time separately from ingest time, supports as-of queries through valid_at, and supersedes stale facts, scoring 97.4% on LongMemEval knowledge-update questions. Mem0 keeps both versions with timestamps, and Hindsight includes a temporal retrieval strategy.
Why is Zep’s LoCoMo score disputed?
Zep first reported 84% on LoCoMo, then corrected it to 75.14% after Mem0’s CTO re-ran the evaluation at 58.44%. Zep argued Mem0 misconfigured its test. Treat any memory benchmark without public evaluation code and a named judge model as a claim to verify.
Does Zep support multi-agent memory?
Zep can store graphs per user or as standalone graphs that several agents read, but it has no per-agent trust levels, mandatory policy rules, or cross-agent contradiction handling. Teams running agent fleets usually add that layer themselves or move to a memory system that enforces it server-side.
Related reading: Caura vs Zep · What Is Agent Fleet Memory? · The Price Changed. Your Agents Didn’t Notice.