Certified benchmark · LoCoMo · July 2026

The number, and how to cite it.

Papez scores 85.55 ± 0.37 on LoCoMo under a frozen protocol. This page is the citation surface: the exact quotable claim, the ten certified runs behind it, the comparison, and the disclosures, so a publisher or a model can extract one accurate sentence and nothing more.

The citation

One sentence, exact.

Quotable claim, verbatim

Papez scores 85.55 ± 0.37 on LoCoMo, certified across ten runs with a frozen gpt-4o-mini answerer and judge at temperature 0 (n=1,540, categories 1–4; evaluated July 2026).

The exact sentence, copy it verbatim.
BibTeX
@misc{papez-locomo-2026,
  title        = {Papez: LoCoMo Benchmark Results},
  author       = {Astrix Labs},
  year         = {2026},
  month        = {jul},
  note         = {J-score 85.55 ± 0.37, mean of 10 runs;
                  frozen gpt-4o-mini answerer and judge, temperature 0;
                  n=1540, LoCoMo categories 1--4},
  howpublished = {\url{https://papez.ai/benchmarks/locomo}}
}
The protocol

What was frozen.

The answerer and judge are the strictest published configuration, the Mem0 paper's setup (arXiv:2504.19413), reused verbatim by Zep. Whatever is frozen defines comparability.

EvaluatedJuly 2026
DatasetLoCoMo
Questions1,540 (n)
Categories1–4
Answerergpt-4o-mini · frozen
Judgegpt-4o-mini · frozen
Temperature0
Runs10 independent full runs
The certified runs

Ten runs. One number.

Ten independent full runs of the frozen protocol. The certified figure is their mean; the spread is the standard deviation across runs.

RunLoCoMo J-score
0184.87
0285.84
0385.71
0485.58
0585.26
0685.78
0785.52
0886.17
0985.52
1085.19
Mean ± std85.55 ± 0.37
Comparison

Same yardstick.

Published numbers for systems measured under a comparable LoCoMo setup, same dataset, same J-score metric.

SystemLoCoMo score
Papez85.55
Zep75.14
Mem066.9

Comparability note. Cross-vendor LoCoMo numbers are only comparable when the answer model, the judge model, and the question set match. The rows above share the dataset and metric; where a vendor's published pipeline differs, treat small gaps as not meaningful. Marketing figures above 90 that use non-comparable configurations are excluded here.

Non-comparable, disclosed: swapping only the answering model for a current one (gpt-4.1-mini), same memories, same retrieval, same judge, scores 89.68. It is never presented as the headline, because it changes the frozen answerer and so is not comparable to the numbers above.

Limitations & disclosures

What the number does not say.

Every figure that qualifies the headline, published with it, not buried. If a memory vendor tells you their system only adds correct answers, they have not measured.

CeilingAn oracle probe, answering from gold evidence alone, measures the ceiling of this frozen protocol at 94.9. Beyond it lie defective gold labels, judge strictness, and the answerer's own reasoning limits.
ChurnAt temperature 0, re-running the identical configuration still regenerates 5.4% of answers differently. The ± 0.37 spread across ten runs is the honest measure of that noise.
Gate G2-PThe certified configuration persistently answers 170 questions the naive baseline cannot, and persistently loses 72 it got right, a 2.4 : 1 trade. Our persistence gate flags this as exceeding the harm budget: RED. Certified as net-better with a disclosed trade, not as strictly better.
Non-comparableA run with a modern answerer (gpt-4.1-mini) scores 89.68. It changes the frozen answerer, so it is disclosed as a non-comparable row and never used as the headline.
Verify it

Check our work.

The full method, frozen protocol, per-category results, the failure ledger, and the measured ceiling, is the dossier. The engine is open source (AGPLv3). The raw run figures on this page are published as JSON.