Evidence registry

Current benchmark evidence

Only active claims with approved wording appear below. Withdrawn, control, and withheld records remain available in the machine-readable registry for auditability.

Caura retrieval-augmented LoCoMo accuracy

Active

Caura scored 77.9% (1,199/1,540) under its documented LoCoMo semantic-judge protocol using the retrieval-augmented agentic-v1 pipeline.

All 1,540 scored LoCoMo questions in categories 1-4.

Answering model
Google Gemini gemini-3.8-flash
Judge
Google Gemini gemini-3.8-flash

Caveats

  • This is the Caura retrieval-augmented result, not the full-context control.
  • The semantic-judge protocol is not the original LoCoMo token-F1 metric.
  • The answering model and semantic judge are both recorded as gemini-3.8-flash; this is not an independent-model judge result.
  • The repository documents the summary but does not currently commit the full question-level result artifact.

No raw result artifact published.

LongMemEval accuracy under the reference judge

Active

Caura answered 461 of 500 LongMemEval_S questions correctly (92.2%) under the benchmark's GPT-4o reference judge.

All 500 LongMemEval_S questions with one configuration across all six question types.

Answering model
Google Gemini gemini-3.8-flash
Judge
OpenAI gpt-4o

Caveats

  • Reader prompts were iterated against failures on this benchmark; this is not a held-out evaluation.

LongMemEval accuracy under the secondary judge

Active

The same 500 frozen LongMemEval_S answers scored 90.2% (451/500) under the secondary Gemini 3.5 Flash-Lite judge.

The same 500 saved answers as the reference-judge result.

Answering model
Google Gemini gemini-3.8-flash
Judge
Google Gemini gemini-3.5-flash-lite

Caveats

  • This is a robustness result and must not replace the reference-judge headline.

LongMemEval compact-context token savings

Active

On LongMemEval_S, the median compact retrieved context was 22,410 tokens versus a 107,706-token full haystack: 79.2% context-only savings; counting every reader call yields 75.4%.

Median token counts across all 500 headline questions, using the answering model's tokenizer.

Answering model
Google Gemini gemini-3.8-flash
Judge
Not applicable

Caveats

  • 79.2% is context-only; the all-reader-call comparison is 75.4%.

Machine-readable source

Download the exact vendored registry at /evidence/claims.json. This copy is pinned to upstream commit 8b5776412d3c0c07fb4999b7aa3432ff76809b78.