Agent MemoryLettaComparison

6 Best Letta Alternatives for Agent Memory and Runtimes in 2026

Letta archived its Python server for TypeScript Letta Code. Where that leaves your memory — and the six alternatives that each win a narrower job.

October 5, 2026 · Caura.AI

The best Letta alternative for teams whose agents already run in LangGraph, CrewAI, the OpenAI Agents SDK, or their own loop is Caura. It adds shared, governed memory without replacing your runtime, and it captures memory in code around the agent instead of relying on the model to save it.

If you want to replace Letta’s whole runtime instead, pick Mastra for TypeScript or LangGraph with LangMem for Python. That choice matters now because Letta moved its Python server to an archive branch in August 2026 and continues development in Letta Code, a TypeScript agent runtime.

Letta alternatives covered:

  • Caura: governed shared memory that plugs into your existing agents, with code-driven and scheduled capture.
  • Mastra: a TypeScript agent framework with working memory and Observational Memory.
  • LangGraph + LangMem: the Python runtime and memory toolkit closest to Letta’s self-editing pattern.
  • Mem0: a drop-in memory API for a single agent.
  • Zep: a temporal knowledge graph for facts that change over time.
  • Hindsight: a self-hosted memory engine with background reflection, under MIT.

Why do teams choose Letta in the first place?

Letta started as MemGPT, the paper that treated the context window like RAM and let the model page information in and out of long-term storage. That idea still defines it. A Letta agent has core memory blocks that sit in its prompt, an archival store it searches, and built-in tools such as memory_insert and memory_replace that let it rewrite its own memory as it works.

Teams pick Letta for three reasons. Agents are persistent services with identity, so the same agent can run for months, and you can inspect exactly what is in its context. Sleep-time agents consolidate memory in the background between conversations. And in Letta Code, all context is tracked in git through MemFS, which makes memory changes reviewable. The letta-ai/letta repository has 25.0k GitHub stars as of 30 September 2026.

Those are real strengths. The search for Letta alternatives usually starts with what surrounds them.

Why are Letta users looking for alternatives?

Six problems come up in Letta’s own repositories, docs, and pricing. Several are recent enough that most comparison pages have not caught up.

1. The Python server was archived, and development moved to TypeScript

The letta-ai/letta README now says the current source lives in letta-ai/letta-code, an npm package, and that the retired Letta V1 API server sits on an archive branch. The archive commit landed on 16 August 2026. A month earlier, the README had already pointed developers to the TypeScript Letta Agent SDK, and in September security reports were limited to maintained projects. Letta’s docs list Docker self-hosting under “Deprecated.”

Teams that built a Python service on the Letta server now face a migration either way.

The practical effect shows in the issue tracker. Requests like conversation import for agent migration and archival memory deduplication were closed as not planned.

2. You adopt a runtime, and leaving it means rewriting the agent

Letta runs the agent loop, tool execution, and state. If your agents already live in another framework, Letta replaces that stack instead of sitting beside it. Leaving is harder still: agent files export configuration, memory blocks, and tools, but issue #3237 shows there is no way to create a conversation with its existing messages, so history is lost when an agent moves between projects.

Cost scales with the runtime too. As of 30 September 2026, the Letta API plan is $20 a month plus $0.10 per active agent, $0.00015 per second of tool execution, and pass-through LLM usage. Role-based access control and SSO sit on Enterprise.

3. Memory quality depends on the model remembering to save

Self-editing memory means the model decides, mid-task, whether to write something down. When it is busy, it skips the write and nothing errors. In a read-only analysis of an eToro deployment snapshot dated 25 August 2026, Caura measured what agents kept when left to themselves: 96% of what they logged was disposable telemetry, recalled essentially never.

Tool-driven memory also raises the bar for models. Local setups repeatedly broke on it, with issues such as Letta 0.9.1 not working with the latest Ollama and agent creation failing with Ollama as provider.

4. Compaction can erase the conversation

When context fills, Letta summarizes older messages. In issue #3270, a sliding-window setting meant to evict 15% of messages wiped the entire history and left one summary, and the issue was closed as not planned. Issue #3279 shows compaction failing outright when a global context cap also limited the summarizer model.

5. Archival memory fills with duplicates

Core memory has rewrite tools. Archival memory has no equivalent consolidation. The request in issue #3116 shows four passages all saying the user likes blue, and it was closed as not planned. Every duplicate is another result the agent has to read past on each search.

6. Agents can act on themselves and on their own memory

Letta Code agents with shell access deleted themselves in production twice, and the reporter noted that client-side deny rules could be bypassed with base64-encoded commands. The reporter proposed a server-side delete_protected flag; the issue later closed as stale. The same openness applies to memory: text an agent reads can be written into its own core memory and then sit in its prompt on every future turn. Memory injection research shows attackers can plant records through ordinary queries alone, which makes model-controlled writes a governance question as well as a quality one.

Where is Letta actually strong, and what fully replaces it there?

Letta does five jobs at once, and a replacement only counts as complete if it covers the jobs you actually use.

The runtime and self-edited working memory

This is Letta’s signature, and two options replace it whole. Mastra gives TypeScript teams an agent framework with working memory the agent updates and Observational Memory for long histories. LangGraph with LangMem gives Python teams a stateful runtime plus hot-path memory tools that work like Letta’s self-editing, and a background manager that consolidates. If Letta’s archive left a Python service stranded, LangGraph is the shortest path that keeps you in Python.

Long-lived memory and background consolidation

If what you valued was an agent that remembers for months and cleans up between sessions, you can keep your own runtime. Caura’s crystallizer merges near-duplicate memories and its Interviewer writes memories from transcripts on a schedule. Hindsight’s reflect operation builds observations and mental models from stored facts.

Memory that many agents share

Letta can share memory blocks and let agents search each other’s conversations, but it has no per-agent trust levels, mandatory policies or cross-agent contradiction handling. That job is where Caura is built to lead.

How do the best Letta alternatives compare?

The table scores the six Letta alternatives on the jobs above and on how they write memory.

ToolReplaces runtimeWho writes memoryConsolidationMulti-agent governanceLicense
CauraNo, plugs into yoursModel, code or scheduleCrystallizer, supersessionTrust tiers, keystones, audit logApache 2.0
MastraYesFramework and modelObservational MemoryNone built inApache 2.0 core
LangGraph + LangMemYesModel tools or background managerBackground managerNone built inMIT
Mem0NoApplication calls add()Pro plan onlyAgent ID scopingApache 2.0
Zep (Graphiti)NoApplication sends episodesFact invalidationNone built inApache 2.0
HindsightNoApplication calls retainReflect, observationsPer-bank isolationMIT

1. Caura: shared memory for the agents you already run

Repository: github.com/caura-ai/caura

Caura is the Letta alternative for teams that want Letta’s long-lived, self-improving memory without moving their agents into Letta’s runtime. It connects over MCP to Claude Code, Cursor, and any MCP client, and has integration guides for LangChain, LlamaIndex, CrewAI, AutoGen, and the OpenAI Agents SDK, listed on the MCP server page. Your agent keeps its loop. Caura keeps the memory.

How Caura answers each reason teams leave Letta

No runtime migration: Because Caura is a memory layer, an archived server or a language switch in your agent framework does not strand your memory. It lives in PostgreSQL with pgvector, which you can self-host from the open-source release with docker compose up.

Memory the model cannot forget to write: Caura supports three capture modes. Agents can still write through tools, as in Letta. Deterministic capture runs recall before each turn and commit after it, as code, through the OpenClaw plugin or the Rail SDK preview. And the Interviewer reads transcripts on a schedule and writes typed memories. In Caura’s read-only analysis of an eToro deployment snapshot dated 25 August 2026, it recovered 48% of decisions and 66% of preferences, and those memories were four times more likely to be recalled by an agent other than their author. The case against agent journaling explains why this path exists.

Consolidation built in: The memory crystallizer merges small memories about the same entity into denser ones, and contradiction detection supersedes stale facts instead of storing both. That is the archival cleanup requested in Letta issue #3116.

Permissions are enforced on the server. Every call checks the agent’s trust level. Standard agents read and write only in their own fleet, and only admin-level agents can delete memories, so a standard agent’s delete returns 403, whatever command it runs. New agents can be set to start restricted until an operator approves them. Keystones deliver mandatory rules to every agent at session start. Every write and delete is audit-logged.

Pricing that ignores agent count. Every plan, including Free, allows unlimited agents. As of 30 September 2026, Pro is $49 a month ($41 billed annually) on the pricing page.

If your agents already run somewhere other than Letta and you only adopted Letta for memory, Caura’s Letta comparison lists each capability side by side.

Caura benchmark results and production proof

Caura scores 92.2% on LongMemEval, 461 of 500 questions under the GPT-4o reference judge, with a 22.4k-token median context. In a separate warm, single-tenant benchmark, search ran at 23 ms p50. Letta does not publish a LongMemEval score. The evaluation code is public in caura-longmemeval. In production, eToro’s Company Brain runs 300+ agents on Caura with 26,500+ memories and 1,372 shared skills.

Where Caura is weaker: it is not an agent runtime, so you get no agent loop, scheduling or Letta-style ADE. Its community is far smaller than Letta’s.

Best for: teams with agents in any framework who want persistent, shared memory, and teams that need permissions and an audit trail before agents can write to shared state.

2. Mastra: the TypeScript runtime replacement

Repository: github.com/mastra-ai/mastra

Mastra is the closest match if you liked Letta’s all-in-one design and are already moving to TypeScript, which is where Letta itself went. It bundles agents, workflows, RAG, evals and memory. Its memory has conversation history, working memory the agent maintains, and Observational Memory, which compresses long histories into observations.

Mastra publishes a 94.87% LongMemEval result for Observational Memory with gpt-5-mini, which is vendor-reported. Check which features sit in ee/ directories before production use.

Best for: TypeScript teams replacing Letta end to end with one framework.

3. LangGraph with LangMem: the Python runtime replacement

Repositories: github.com/langchain-ai/langgraph and github.com/langchain-ai/langmem

LangGraph supplies the stateful runtime: checkpointed graphs and a long-term memory store. LangMem supplies memory management tools the agent calls “in the hot path,” which is the same pattern as Letta’s self-editing, plus a background memory manager that extracts, consolidates and updates knowledge after the fact. Together they keep Python teams in Python. LangMem is far smaller than LangGraph, so expect to read source when docs run out.

Best for: Python teams whose Letta service was stranded by the archive and who want the self-editing pattern without Letta.

4. Mem0: drop-in memory for a single agent

Repository: github.com/mem0ai/mem0

Mem0 replaces Letta’s archival memory with two calls, add() and search(), and leaves your runtime alone. Your application decides when to write, which removes the model-must-remember problem. Its v3 algorithm keeps both old and new versions of a changed fact, so current-state resolution falls to your code, as users describe in issue #4956. As of 30 September 2026, graph memory and consolidation are on the $249 Pro plan.

Best for: single-agent personalization where Letta was more framework than you needed.

5. Zep: temporal memory for facts that change

Repository: github.com/getzep/graphiti

Zep stores facts as edges with validity windows in its Graphiti engine, so it can answer what a user’s plan was in March as well as today. That is a precise upgrade over Letta’s text memory blocks when your agent tracks changing account or health data. Zep Cloud bills per byte written, and self-hosting now means running Graphiti with your own graph database. Our Zep comparison covers those trade-offs in full.

Best for: single-agent products where time-aware recall is the core requirement.

6. Hindsight: self-hosted memory that reflects

Repository: github.com/vectorize-io/hindsight

Hindsight covers the part of Letta that sleep-time agents handle. It organizes memory into banks with three operations: retain to store, recall to search, and reflect to build observations and mental models from what it holds. It runs as one Docker container with embedded Postgres, and its README states that its LongMemEval results were independently reproduced by Virginia Tech’s Sanghani Center.

Best for: teams that want background consolidation and strong recall under MIT, on their own infrastructure.

How do you choose the right Letta alternative?

Start from the reason you are leaving Letta, then answer four questions before you migrate.

Do you want to keep your agent framework?

If your agents already run in LangGraph, CrewAI or a custom loop, pick a memory layer: Caura, Mem0, Zep or Hindsight. If you want one framework to own everything the way Letta did, pick Mastra or LangGraph with LangMem.

Python or TypeScript?

Letta’s move to Letta Code makes this the first fork. Mastra is TypeScript-only. LangGraph and LangMem are Python-first. Caura, Mem0, Zep and Hindsight work from either through REST, MCP or SDKs.

Should the model or the code decide what gets remembered?

Letta, Mastra’s working memory and LangMem’s hot-path tools let the model decide. Mem0, Zep and Hindsight let your application decide. Caura supports all three and adds scheduled reflection. Caura’s write-up on deterministic memory compares what Letta, LangGraph, Zep and Mem0 each ship for this.

Will more than one agent write to the same memory?

Letta Code agents can call each other as subagents, but shared writes still need to know who wrote what and who may read it. Caura’s guide to how multi-agent systems coordinate covers the collisions that appear once they do.

Which Letta alternative should you pick?

Pick Caura if your agents run outside Letta, if more than one agent shares memory, or if you need memory captured by code instead of by the model’s discretion. It is the only option here with code-driven and scheduled capture, per-agent trust levels and an audit trail on every write, and it runs in production at 300+ agents.

Pick Mastra to replace Letta end to end in TypeScript, or LangGraph with LangMem to do it in Python. For a single agent, pick Mem0 for the simplest API, Zep for facts that change over time, and Hindsight for self-hosted consolidation under MIT.

If you are on Letta today, the quickest test is to connect one of your existing agents to Caura’s free tier over MCP, run it for a day without adding any save calls to your prompt, and look at what the Interviewer recorded. That shows how much of your agent’s work Letta’s model-driven writes were leaving behind.

Frequently Asked Questions

Is Letta still maintained?

Yes, as Letta Code. The original Python API server moved to an archive branch of letta-ai/letta in August 2026, and active development happens in letta-ai/letta-code, a TypeScript agent runtime installed from npm. Existing tags remain available, but new projects are directed to Letta Code and the TypeScript Letta Agent SDK.

What is the difference between Letta and MemGPT?

MemGPT was the 2023 research paper and project that introduced OS-style memory paging for LLMs. It was renamed Letta in 2024 and grew into a full agent platform with memory blocks, an agent development environment and, now, the Letta Code runtime.

What is the best open-source Letta alternative?

For shared memory across agents, Caura under Apache 2.0. To replace the runtime, LangGraph with LangMem under MIT for Python, or Mastra with an Apache 2.0 core for TypeScript. Hindsight is the strongest MIT-licensed memory engine for a single agent.

Can I migrate my Letta agents without losing history?

Partly. Letta agent files export configuration, memory blocks and tools, but there is no API to recreate a conversation with its past messages, and that request was closed as not planned. Export messages through the conversations API and write them into your new memory layer as a batch before you switch.

Does Letta support multi-agent memory?

Letta can attach shared memory blocks to several agents, and Letta Code agents can search each other’s conversations and call each other as subagents. It does not provide per-agent trust levels, mandatory policy rules or cross-agent contradiction handling, which teams running fleets usually need.

How much does Letta cost?

As of 30 September 2026, Letta is free for up to 3 stateful agents with your own API keys, and Pro is $20 a month for up to 20. The developer API plan is $20 a month plus $0.10 per active agent, $0.00015 per second of tool execution and LLM usage at token cost. RBAC and SSO are on Enterprise.

Related reading: Caura vs Letta · What Is Agent Fleet Memory? · Reflective Memory: The Interviewer