Benchmarks · KnowMe-Bench

66.59%~43–48%

Substrate 66.59 percent versus Mem0 paper-official results of roughly 43 to 48 percent, depending on backbone, on KnowMe-Bench.

SubstrateMem0 · paper-official

+18.6 to +23.8 points for Substrate

Mem0's published KnowMe-Bench results aggregate to the low-to-mid forties. When we re-ran Mem0 OSS ourselves on a stronger backbone than its paper used, it climbed to 65.09 — and Substrate still leads. One full pass: 2,580 queries, same answer model, same judge.

The scorecard

One benchmark. Every number traceable.

Hover any figure for the receipt behind it — where the number comes from and how it was produced.

KnowMe-Bench aggregate results: Substrate versus Mem0 — our controlled re-run and Mem0's paper-official runs
SystemAggregate scoreQueriesAnswer model & judge
SubstrateOur run66.59%ReceiptOne full 2,580-query aggregate run of KnowMe-Bench. Method and analysis: the KnowMe-Bench paper.2,580ReceiptEvery query in the benchmark, one full pass. No sampling, no subset.Same model, same judge
Mem0 OSS (strengthened)Our re-run65.09%ReceiptOpen-source Mem0, re-run by us on Minimax — a stronger answer backbone than either backbone in the paper's runs. Same queries, answer model, and judge as our Substrate run.2,580ReceiptIdentical query set as the Substrate run.Same model, same judge
Mem0 (paper)Paper run · Qwen3-32B~42.8%ReceiptSeven-task aggregate for Mem0 with the Qwen3-32B backbone, computed from Table 1 of KnowMe-Bench.7 tasksQwen3-32B · paper setup
Mem0 (paper)Paper run · GPT-5-mini~48.0%ReceiptSeven-task aggregate for Mem0 with the GPT-5-mini backbone, computed from Table 1 of KnowMe-Bench.7 tasksGPT-5-mini · paper setup
vs paper-official Mem0+18.6 to +23.8 points—Range across both paper backbones
vs strengthened re-run+1.50 points—Controlled comparison

The paper runs used different backbones than our re-run; the paper aggregates are computed from Table 1 of KnowMe-Bench.

Benchmark
KnowMe-Bench — citation-backed memory recall
Our runs
Full 2,580-query aggregate passes, both systems
Paper baseline
Published Mem0 results — seven-task aggregates from Table 1
Artifact
Aggregate scores and method notes, downloadable below

Methodology

Check the work yourself.

How we ran it

We ran KnowMe-Bench end to end. All 2,580 queries, one full aggregate pass, both systems on the same footing: the same answer model produces every response, and the same judge grades it against the published rubric.

In the controlled re-run we hold the answer model and judge constant, so the gap between Substrate and the strengthened Mem0 baseline measures what actually differs between the systems — the memory layer. And that baseline is no soft target: it gives Mem0 OSS a stronger backbone than its own paper runs had. The benchmark itself is public: read the paper, pin the exact upstream commit we evaluated against, and grade with the published rubric.

The downloadable artifact contains the aggregate scores and method notes for both systems, exactly as run.

Inspect the sentence

Fluent is not the same as true.

Substrate publishes a clean summary without hiding how it was made. Evidence, time, uncertainty, and corrections stay attached to the memory.

Cited

Every published sentence keeps a route back to the sources that support it.

Dated

Observed and effective dates keep an old fact from silently appearing current.

Qualified

Current, historical, superseded, inferred, conflicting, and uncertain context remain distinct.

Correctable

An edit becomes a durable fact. The earlier version and its evidence remain traceable.

Example

One sentence. Three things you can inspect.

“Acme Corp's renewal is at risk.”

  • SourceGmail · Dana Reyes · Jul 28
  • ConfidenceHigh · direct statement
  • StatusCurrent · not committed for Q4

Start

Build memory you can inspect.

Open beta. One plan at $20 a month.

Login / Sign upNow in open beta