Caura retrieval-augmented LoCoMo accuracy
ActiveCaura scored 77.9% (1,199/1,540) under its documented LoCoMo semantic-judge protocol using the retrieval-augmented agentic-v1 pipeline.
All 1,540 scored LoCoMo questions in categories 1-4.
- Answering model
- Google Gemini gemini-3.8-flash
- Judge
- Google Gemini gemini-3.8-flash
Caveats
- This is the Caura retrieval-augmented result, not the full-context control.
- The semantic-judge protocol is not the original LoCoMo token-F1 metric.
- The answering model and semantic judge are both recorded as gemini-3.8-flash; this is not an independent-model judge result.
- The repository documents the summary but does not currently commit the full question-level result artifact.
No raw result artifact published.