hybrid recall@1
83.2% recall@1 and 98.8% recall@5 on the public split.
500 public questions and 23,867 ingested sessions; offline, judge-free retrieval. Not end-to-end QA, customer performance, or production validation.Read the series report ↗Results and methods
These are company-run reference measurements. Each card names the test, the scope, and the limit.
83.2% recall@1 and 98.8% recall@5 on the public split.
500 public questions and 23,867 ingested sessions; offline, judge-free retrieval. Not end-to-end QA, customer performance, or production validation.Read the series report ↗In the tested Perseus Vault configuration.
1,378 / 3,750 blank or abstaining (36.7%); wrong counted separately. Self-run comparison, not operational validation.Read the full comparison ↗Company-run retrieval on the 500-instance LongMemEval split.
Retrieval only. No answer-quality control arm.Read the retrieval report ↗Naive full-file assembly versus the shipped-default render on a named repo corpus.
Cold cache, context + 14 prompts, five fixture documents. Content assembly only; not provider billing or dollar savings.Read the token report ↗content-hashed scale report at 100K entities on named hardware.
Throughput only. Not customer capacity, deployment performance, or an SLA.Read the scale report ↗Deprecated answerer/judge experiments remain internal-only and are excluded from the current public metric.
Read before using
They do not establish a mission result. They show what was measured in named company-run tests.
No ATO/cATO, CMMC certification, classified-data authority, or mission-system authority is claimed.
No customer, operational deployment, or Government authorization is represented by these measurements.
Weighted token workload and context reduction are not provider billing, cost savings, or a service-level promise.
Detailed reports