1,986 questions across all ten conversations in under eight minutes, for about $3. The whole benchmark loaded into Caura in 47 seconds, zero errors.
LoCoMo is the benchmark people reach for when they want to know whether an AI agent can remember. Ten long conversations between two people, each spread over months of dated sessions, and nearly two thousand questions whose answers are buried somewhere in the history. What did she paint last summer? When did he go to the support group? How long ago was the birthday?
It is also the benchmark people most often run in ways nobody else can check. Scores get published without the prompt that produced them. Harnesses take hours, so nobody re-runs them. Whole categories get quietly dropped. The number travels; the method doesn’t.
We wanted the opposite: the entire benchmark, fast enough to run on a whim, with every result checkable by anyone. So we built it, ran it, and published it at caura-ai/caura-locomo.
What LoCoMo measures
Each of the ten conversations runs to 19–32 dated sessions. Every question comes tagged with the dialogue turns that hold its answer, which is what makes the benchmark scorable automatically. The questions fall into five types:
| type | questions | asks for |
|---|---|---|
| single-hop | 841 | one fact from one turn |
| multi-hop | 282 | facts combined across sessions |
| temporal | 321 | dates and durations |
| open-domain | 96 | inference beyond what was said |
| adversarial | 446 | things that were never mentioned — the right answer is to say so |
The whole benchmark, in the time it takes to make coffee
A full LoCoMo pass through caura-locomo — all 1,986 questions, answered and judged — finishes in under eight minutes and costs about three dollars.
Two design choices do most of the work. Questions run in parallel instead of one at a time. And every prompt puts the conversation first and the question last, so the provider caches the shared history across the two hundred or so questions about each conversation: on our headline run, 97.7% of the reader’s input tokens were served from cache.
Loading the benchmark into Caura is just as quick. The harness writes each conversation as one memory per session and one agent per conversation — 272 sessions across all ten conversations in 47 seconds, with zero errors, using Caura’s bulk write in strong mode so every memory is searchable the moment the write returns. Retrieval sweeps call no language model at all, so they cost nothing.
A benchmark that takes a day to run gets run once and quoted forever. One that takes eight minutes gets run every time something changes — and a claim you can re-run is a claim you can check.
Three things the usual LoCoMo harnesses get wrong
Running a benchmark end to end, quickly and many times over, surfaces things a single slow run never will. These three are fixed in caura-locomo.
The adversarial category has ground truth — all 446 questions of it
LoCoMo’s adversarial questions ask about things that were never said, and the right answer is to decline. Published evaluations, mem0’s among them, drop the category on the grounds that its answers are missing. They aren’t: every one of the 446 items carries an adversarial_answer field. caura-locomo scores the category properly, offering every question the chance to decline so the test measures a real decision rather than the prompt. Scored that way, the reader turned down 76% of the unanswerable questions while wrongly declining only 6.5% of the answerable ones.
The category labels ship scrambled
LoCoMo numbers its question types, and harnesses we found in circulation map those numbers to the wrong names — so per-category results get reported under the wrong heading. Single-hop is category 4, the 841-question bucket, not category 1. caura-locomo pins the correct mapping against the dataset with a test that fails if it ever drifts.
The prompt is worth 24 points
Same reader, same judge, same questions: three rounds of prompt rewrites moved the score by 24 points. That is the single biggest lever on a LoCoMo number, and it is almost never published alongside one. caura-locomo writes the exact prompt text and its hash into every result file, so the prompt travels with the score.
We also built in something retrieval results badly need: a line to compare them against. With around 27 sessions per conversation, fetching 50 of them returns nearly the whole conversation by construction. locomo baseline prints what random selection would score at every depth, so a retrieval number can be read for what it is.
caura-locomo: every number, reproducible
Everything behind this post is at github.com/caura-ai/caura-locomo, MIT-licensed. The harness runs three arms — full conversation, retrieval, and a free retrieval-only sweep — reports two judges on every run, and ships the committed results for the figures above. We tested it the way a newcomer would: a clean clone, a fresh install, and the published figures reproduced.
- Install.
uv sync— one command, test suite included. - Fetch the benchmark.
uv run locomo download, thenuv run locomo verifyto check your copy against the official file. - Run it.
uv run locomo run --name mine --arm fullcontext --no-derive— the whole benchmark, in minutes. - Judge it twice.
uv run locomo rejudge outputs/mine/results.json --judge-model gpt-4oputs the headline judge beside the fast one. - Bring your own memory. Point the retrieval arm at a store and read it against the random line.
Benchmarks you can run are benchmarks you can trust
This is the second benchmark we have opened up this month, after caura-longmemeval. The idea behind both is the same. Memory for AI agents is going to be judged on numbers like these, and a number is only as good as the next person’s ability to reproduce it. We would rather compete on benchmarks anyone can run than on scores nobody can check.
So the method is public and the runner is public. Clone it, run it, and tell us what you find.
Related reading: Caura on LongMemEval · Caura benchmarks · What is AI agent memory?