Results and methods

Numbers first. Conditions beside them.

These are company-run reference measurements. Each card names the test, the scope, and the limit.

Conflict handling / MemConflict18

wrong answers out of 3,750 questions

In the tested Perseus Vault configuration.

1,378 / 3,750 blank or abstaining (36.7%); wrong counted separately. Self-run comparison, not operational validation.Read the full comparison ↗
Retrieval99.8%

session recall@10

Company-run retrieval on the 500-instance LongMemEval split.

Retrieval only. No answer-quality control arm.Read the retrieval report ↗
Efficiency / Token A/B52.63%

fewer full-request prompt tokens

Naive full-file assembly versus the shipped-default render on a named repo corpus.

Cold cache, context + 14 prompts, five fixture documents. Content assembly only; not provider billing or dollar savings.Read the token report ↗
Operations40 writes/s

durable write throughput

content-hashed scale report at 100K entities on named hardware.

Throughput only. Not customer capacity, deployment performance, or an SLA.Read the scale report ↗

Deprecated answerer/judge experiments remain internal-only and are excluded from the current public metric.

Read before using

What these numbers do not say.

They do not establish a mission result. They show what was measured in named company-run tests.

Not a certification

No ATO/cATO, CMMC certification, classified-data authority, or mission-system authority is claimed.

Not a customer outcome

No customer, operational deployment, or Government authorization is represented by these measurements.

Not a dollar claim

Weighted token workload and context reduction are not provider billing, cost savings, or a service-level promise.

Reproduce

Start with the public source.

The summary is sanitized. Raw prompts, conversations, credentials, and authorization headers are not published.

Open the Vault repository ↗