Every number we publish is generated by an open harness, committed to this repo, from a public dataset. You don't have to trust us β this page shows how to make the computer tell you the truth yourself. Three budgets, pick one:
| You have... | Path | Time | Cost |
|---|---|---|---|
| 5 minutes | Verify the committed results | ~1 min | $0 |
| ~1 hour | Run a real subset locally | ~1-2 h | $0 |
| ~$10 | Full 500-question run | ~3 h | ~$5-10 Modal credits |
What you need: a machine with cargo (Rust 1.85+), Python 3.10+, and ~4 GB free
disk. The embedding model (~200 MB) downloads once on first run. Everything else is
in this repo.
Published numbers (v0.19.0, 500 questions, LongMemEval-S): recall_any@5 98.2% Β· recall_any@10 98.8% Β· strict recall_all@5 88.0% Β· coverage@5 94.3% β all on the full 500-question basis. On the 470 non-abstention questions (the 30
_absquestions are reported separately: they measure answering, not retrieval), strict recall_all@5 is 88.3% and coverage@5 is 94.6%. Per-release any@5: v0.16.0 98.2% Β· v0.17.0 98.4% Β· v0.19.0 98.2% β the strict family (88.0/95.4) is unchanged across all three. What these metric names mean: docs/benchmarks.md andRESULTS.md.
Committed raw outputs (verify without running anything β section 1 below):
| File | Run | Headline |
|---|---|---|
results/default-500q-v0.19.0.jsonl |
v0.19.0 (current release binary) | any@5 98.2% |
results/default-500q-v017.jsonl |
v0.17.0 | any@5 98.4% |
results/default-500q.jsonl |
v0.16.0 (canonical validation) | any@5 98.2% |
The raw per-question output of the full 500-question runs is committed under
results/ β one JSONL per measured release (table above). The harness
stores the retrieved session ids (ranked) for every question, and the dataset
stores which sessions are correct (answer_session_ids). Scoring is a 20-line
join β swap the filename to check either release:
cd benchmarks/longmemeval
python3 - <<'EOF'
import json
RAW = 'results/default-500q-v0.19.0.jsonl' # v0.19.0 (headline); or default-500q-v017.jsonl / default-500q.jsonl for older releases
rows = {}
for line in open(RAW):
line = line.strip()
if not line:
continue
r = json.loads(line)
rows[r['question_id']] = r['retrieval_results']['retrieved_session_ids']
ds = json.load(open('data/longmemeval_s_cleaned.json'))
gold = {q['question_id']: set(q['answer_session_ids']) for q in ds}
non_abs = [q for q in rows if '_abs' not in q]
any5 = sum(1 for q in non_abs if gold[q] & set(rows[q][:5]))
all5 = sum(1 for q in non_abs if len(gold[q] & set(rows[q][:5])) == len(gold[q]))
cov5 = sum(len(gold[q] & set(rows[q][:5])) / len(gold[q]) for q in non_abs)
print(f"questions scored (non-abstention): {len(non_abs)}")
print(f"recall_any@5 = {any5}/{len(non_abs)} = {any5/len(non_abs):.4f}")
print(f"recall_all@5 = {all5}/{len(non_abs)} = {all5/len(non_abs):.4f}")
print(f"coverage R@5 = {cov5/len(non_abs):.4f}")
print(f"total-miss@50 = {sum(1 for q in non_abs if not (gold[q] & set(rows[q][:50])))}")
EOFExpected output β recall_any@5 = 0.9820 (v0.19.0 file), 0.9840 (v0.17.0)
or 0.9820 (v0.16.0 file), total-miss@50 = 0: every question has all its gold
sessions somewhere in the top-50; everything below perfect is ranking order, not missing
data.
If the file ever stops matching the README, that's a bug β open an issue.
This actually runs uteke: builds the store from each question's haystack, embeds ~2,400 sessions, and measures what the retriever brings back. Deterministic on the same architecture β your numbers should match ours.
cd benchmarks/longmemeval
# 1. Dataset (oracle subset = the 50 hardest-with-known-gold questions, ~5 MB)
./scripts/download_data.sh
# 2. Python deps (stdlib is almost enough; print_metrics.py uses numpy if present)
pip install -r scripts/requirements.txt
# 3. Build uteke from this repo (the harness resolves THIS binary first β
# never a uteke from your PATH)
cd ../.. && cargo build --release -p uteke-cli && cd benchmarks/longmemeval
# 4. Run 50 questions (first run downloads the ~200 MB embedding model once)
python3 scripts/run_eval.py \
--data data/longmemeval_oracle.json \
--output results_mine \
--limit 50
# 5. Score it
python3 scripts/print_metrics.py results_mine/retrieval_results.jsonlNotes:
- Runs pinned to 2 cores at low priority β your machine stays usable.
- ~1-2 h for 50 questions on a modern laptop (embedding is the cost).
- Every question is stored with its own isolated store; no cross-contamination.
- Want the exact published split instead of oracle? Swap in the full dataset:
curl -L -o data/longmemeval_s_cleaned.json \ https://huggingface.co/datasets/xiaowu0162/longmemeval-cleaned/resolve/main/longmemeval_s_cleaned.jsonthen run with--limit 100(or more) on that file.
The published numbers come from the fan-out harness on Modal (10 shards Γ 50 questions, resumed automatically on preemption). Same harness, bigger loop:
cd benchmarks/longmemeval
# dataset (full S split, ~50 MB)
curl -L -o data/longmemeval_s_cleaned.json \
https://huggingface.co/datasets/xiaowu0162/longmemeval-cleaned/resolve/main/longmemeval_s_cleaned.json
# pin the EXACT binary you want to measure (builds from source inside the image;
# do not skip this β an untagged image silently measures whatever it was built with)
export UTEKE_GIT_REF=<commit-sha-or-tag> # e.g. the v0.19.0 release tag
export MODAL_TOKEN_ID=... MODAL_TOKEN_SECRET=...
modal run scripts/modal_fanout.py --strategy default --num-shards 10
# watch progress (shards appear at completion), then score the merged output:
python3 scripts/print_metrics.py results_v017_revalidation/retrieval_results.jsonlTotal cost last run: under $10 (~2.5 h wall clock, 2 preemptions auto-recovered).
No Modal account? The same run_eval.py runs on any Linux box β it just takes
longer serially (~4 min/question on a 4-core desktop).
LongMemEval-S hides facts across ~115 chat sessions per question and asks where the evidence sessions rank. We report two families because they answer different questions:
- recall_any@K β "did the retriever surface the evidence?" (at least one gold session in top-K). This is the metric competitor benchmarks publish; ours: 98.2% @ 5 (full 500-question set, v0.19.0).
- recall_all@K (strict) β "did it get all of them?" (every gold session in top-K; 65% of the 500 questions have multiple gold sessions). The honest ceiling-capable number: 88.0% @ 5 on the full 500 (88.3% on the 470 non-abstention) β mathematical ceiling is 99.4% (3 questions have 6 gold sessions; top-5 physically can't hold all).
- coverage R@K β partial credit (fraction of gold sessions found). 94.3% full-500 (94.6% on the 470 non-abstention).
Per-category breakdown, the 30 abstention questions, contradiction segment, and
the v0.16.0 β v0.17.0 β v0.19.0 comparisons: RESULTS.md,
docs/benchmarks.md.
- Deterministic: same binary + same dataset + same architecture β identical rankings. The harness embeds nothing randomly and stores ranked session ids for every question, so any claim can be re-derived from committed artifacts.
- Pinned provenance: published runs record the exact git ref baked into the
binary (
UTEKE_GIT_REF) and the raw JSONL is committed β no "trust me". - Cross-architecture honesty: ARM vs x86 float noise can swap adjacent ranks on tail questions (documented: 107/108 questions identical in our own re-run, 1 adjacent-rank near-tie with the same top-10 set). Aggregates are stable; individual rank order is stable up to Β±1 near ties.