BenchmarksLoCoMoOpen Source28 September 2026

We ran all of LoCoMo in eight minutes. Then we open-sourced it.

Every question, every conversation, about three dollars — and a runner anyone can clone to check the work.

1,986 questions across all ten conversations in under eight minutes, for about $3. The whole benchmark loaded into Caura in 47 seconds, zero errors.

LoCoMo is the benchmark people reach for when they want to know whether an AI agent can remember. Ten long conversations between two people, each spread over months of dated sessions, and nearly two thousand questions whose answers are buried somewhere in the history. What did she paint last summer? When did he go to the support group? How long ago was the birthday?

It is also the benchmark people most often run in ways nobody else can check. Scores get published without the prompt that produced them. Harnesses take hours, so nobody re-runs them. Whole categories get quietly dropped. The number travels; the method doesn’t.

We wanted the opposite: the entire benchmark, fast enough to run on a whim, with every result checkable by anyone. So we built it, ran it, and published it at caura-ai/caura-locomo.

What LoCoMo measures

Each of the ten conversations runs to 19–32 dated sessions. Every question comes tagged with the dialogue turns that hold its answer, which is what makes the benchmark scorable automatically. The questions fall into five types:

typequestionsasks for
single-hop841one fact from one turn
multi-hop282facts combined across sessions
temporal321dates and durations
open-domain96inference beyond what was said
adversarial446things that were never mentioned — the right answer is to say so

The whole benchmark, in the time it takes to make coffee

A full LoCoMo pass through caura-locomo — all 1,986 questions, answered and judged — finishes in under eight minutes and costs about three dollars.

Two design choices do most of the work. Questions run in parallel instead of one at a time. And every prompt puts the conversation first and the question last, so the provider caches the shared history across the two hundred or so questions about each conversation: on our headline run, 97.7% of the reader’s input tokens were served from cache.

Loading the benchmark into Caura is just as quick. The harness writes each conversation as one memory per session and one agent per conversation — 272 sessions across all ten conversations in 47 seconds, with zero errors, using Caura’s bulk write in strong mode so every memory is searchable the moment the write returns. Retrieval sweeps call no language model at all, so they cost nothing.

A benchmark that takes a day to run gets run once and quoted forever. One that takes eight minutes gets run every time something changes — and a claim you can re-run is a claim you can check.

Three things the usual LoCoMo harnesses get wrong

Running a benchmark end to end, quickly and many times over, surfaces things a single slow run never will. These three are fixed in caura-locomo.

The adversarial category has ground truth — all 446 questions of it

LoCoMo’s adversarial questions ask about things that were never said, and the right answer is to decline. Published evaluations, mem0’s among them, drop the category on the grounds that its answers are missing. They aren’t: every one of the 446 items carries an adversarial_answer field. caura-locomo scores the category properly, offering every question the chance to decline so the test measures a real decision rather than the prompt. Scored that way, the reader turned down 76% of the unanswerable questions while wrongly declining only 6.5% of the answerable ones.

The category labels ship scrambled

LoCoMo numbers its question types, and harnesses we found in circulation map those numbers to the wrong names — so per-category results get reported under the wrong heading. Single-hop is category 4, the 841-question bucket, not category 1. caura-locomo pins the correct mapping against the dataset with a test that fails if it ever drifts.

The prompt is worth 24 points

Same reader, same judge, same questions: three rounds of prompt rewrites moved the score by 24 points. That is the single biggest lever on a LoCoMo number, and it is almost never published alongside one. caura-locomo writes the exact prompt text and its hash into every result file, so the prompt travels with the score.

We also built in something retrieval results badly need: a line to compare them against. With around 27 sessions per conversation, fetching 50 of them returns nearly the whole conversation by construction. locomo baseline prints what random selection would score at every depth, so a retrieval number can be read for what it is.

caura-locomo: every number, reproducible

Everything behind this post is at github.com/caura-ai/caura-locomo, MIT-licensed. The harness runs three arms — full conversation, retrieval, and a free retrieval-only sweep — reports two judges on every run, and ships the committed results for the figures above. We tested it the way a newcomer would: a clean clone, a fresh install, and the published figures reproduced.

  1. Install. uv sync — one command, test suite included.
  2. Fetch the benchmark. uv run locomo download, then uv run locomo verify to check your copy against the official file.
  3. Run it. uv run locomo run --name mine --arm fullcontext --no-derive — the whole benchmark, in minutes.
  4. Judge it twice. uv run locomo rejudge outputs/mine/results.json --judge-model gpt-4o puts the headline judge beside the fast one.
  5. Bring your own memory. Point the retrieval arm at a store and read it against the random line.

Benchmarks you can run are benchmarks you can trust

This is the second benchmark we have opened up this month, after caura-longmemeval. The idea behind both is the same. Memory for AI agents is going to be judged on numbers like these, and a number is only as good as the next person’s ability to reproduce it. We would rather compete on benchmarks anyone can run than on scores nobody can check.

So the method is public and the runner is public. Clone it, run it, and tell us what you find.


Related reading: Caura on LongMemEval · Caura benchmarks · What is AI agent memory?

Frequently asked

What is the LoCoMo benchmark?+

LoCoMo is a long-term conversational memory benchmark: ten long conversations between two people, each spread over 19 to 32 dated sessions, with 1,986 questions whose answers sit somewhere in the history. It covers single-hop recall, multi-hop reasoning across sessions, temporal reasoning, open-domain inference, and adversarial questions about things that were never said.

How long does a full LoCoMo run take with caura-locomo?+

Under eight minutes and about three dollars for all 1,986 questions across all ten conversations. Questions run in parallel, and each prompt puts the conversation first so the shared history is cached across every question about it. Retrieval-only sweeps make no language-model calls and cost nothing.

Does LoCoMo's adversarial category have ground truth?+

Yes. The category is often excluded on the grounds that its answers are unavailable, but all 446 adversarial items carry an adversarial_answer field. The right behaviour is to recognise the question as unanswerable, and caura-locomo scores it by offering every question the chance to decline.

Is caura-locomo open source?+

Yes, MIT-licensed at github.com/caura-ai/caura-locomo. It ships the harness, a random-selection baseline for retrieval, an abstention control, the prompt text and hash in every result file, and the committed runs behind this post. It reproduces them from a fresh clone.