Measured, not claimed
One memory. Two readers.
The same score.
LongMemEval-S, official per-type judge prompts, the real write and recall paths — no benchmark-only shortcuts. A fully local stack and a frontier cloud model land within a point of each other over the same memory. Every number regenerates from the public harness; if you doubt one, run it.
Fully local stack
83%
Local reader
Open-weights 27B reader at temperature 0, served by llama.cpp.
Digestion, embeddings, retrieval, composition and judging all run on local hardware — no cloud service anywhere in the loop.
83 of 100 questions correct, every row graded.
Cross-check
84%
Cloud reader, same memory
Claude Sonnet composes over the identical retrieved statements, in a stripped environment with no tools and no ambient context.
Digestion, retrieval and judging stay local — only the composing model changes.
84 of 100. One point apart: the memory, not the reader, sets the score.
By question type
16-17 questions per type, 100 total, every row graded. Full per-question verdicts in the dated run log.
How to read this honestly
We wrote the rules we ask you to grade benchmarks by. Apply them here first.
- A 100-question stratified slice of LongMemEval-S, not the full 500. Binomial noise at n=100 is roughly ±7 points — treat smaller differences, including the local-vs-cloud gap above, as noise.
- The judge is a local open-weights model at temperature 0 with the official per-type LongMemEval prompts, not the GPT-4o the paper used. A different judge can shift absolute numbers; it applies equally to both readers.
- The cloud reader is not deterministic — repeated passes over identical memory moved 2-3 verdicts per hundred. The local reader and the judge are temperature-0 and reproducible.
- The reader sees only the retrieved statements, none of the identity envelope a live session carries. That isolates what the memory retrieved — and costs points a fuller context would catch.
- Both readers miss the same hard core: counting occurrences across many sessions where the haystack itself tells one event in conflicting ways. That agreement is the result we consider load-bearing.
The five red flags we ask you to check every memory benchmark against — including this one — live on the credibility page.