Skip to content

Latest commit

Β 

History

History
181 lines (141 loc) Β· 8.28 KB

File metadata and controls

181 lines (141 loc) Β· 8.28 KB

Reproducing the LongMemEval-S results β€” run it yourself

Every number we publish is generated by an open harness, committed to this repo, from a public dataset. You don't have to trust us β€” this page shows how to make the computer tell you the truth yourself. Three budgets, pick one:

You have... Path Time Cost
5 minutes Verify the committed results ~1 min $0
~1 hour Run a real subset locally ~1-2 h $0
~$10 Full 500-question run ~3 h ~$5-10 Modal credits

What you need: a machine with cargo (Rust 1.85+), Python 3.10+, and ~4 GB free disk. The embedding model (~200 MB) downloads once on first run. Everything else is in this repo.

Published numbers (v0.19.0, 500 questions, LongMemEval-S): recall_any@5 98.2% Β· recall_any@10 98.8% Β· strict recall_all@5 88.0% Β· coverage@5 94.3% β€” all on the full 500-question basis. On the 470 non-abstention questions (the 30 _abs questions are reported separately: they measure answering, not retrieval), strict recall_all@5 is 88.3% and coverage@5 is 94.6%. Per-release any@5: v0.16.0 98.2% Β· v0.17.0 98.4% Β· v0.19.0 98.2% β€” the strict family (88.0/95.4) is unchanged across all three. What these metric names mean: docs/benchmarks.md and RESULTS.md.

Committed raw outputs (verify without running anything β€” section 1 below):

File Run Headline
results/default-500q-v0.19.0.jsonl v0.19.0 (current release binary) any@5 98.2%
results/default-500q-v017.jsonl v0.17.0 any@5 98.4%
results/default-500q.jsonl v0.16.0 (canonical validation) any@5 98.2%

1) Verify the committed results (5 min, no embedder needed)

The raw per-question output of the full 500-question runs is committed under results/ β€” one JSONL per measured release (table above). The harness stores the retrieved session ids (ranked) for every question, and the dataset stores which sessions are correct (answer_session_ids). Scoring is a 20-line join β€” swap the filename to check either release:

cd benchmarks/longmemeval

python3 - <<'EOF'
import json

RAW = 'results/default-500q-v0.19.0.jsonl'   # v0.19.0 (headline); or default-500q-v017.jsonl / default-500q.jsonl for older releases
rows = {}
for line in open(RAW):
    line = line.strip()
    if not line:
        continue
    r = json.loads(line)
    rows[r['question_id']] = r['retrieval_results']['retrieved_session_ids']

ds = json.load(open('data/longmemeval_s_cleaned.json'))
gold = {q['question_id']: set(q['answer_session_ids']) for q in ds}

non_abs = [q for q in rows if '_abs' not in q]
any5 = sum(1 for q in non_abs if gold[q] & set(rows[q][:5]))
all5 = sum(1 for q in non_abs if len(gold[q] & set(rows[q][:5])) == len(gold[q]))
cov5 = sum(len(gold[q] & set(rows[q][:5])) / len(gold[q]) for q in non_abs)
print(f"questions scored (non-abstention): {len(non_abs)}")
print(f"recall_any@5   = {any5}/{len(non_abs)} = {any5/len(non_abs):.4f}")
print(f"recall_all@5   = {all5}/{len(non_abs)} = {all5/len(non_abs):.4f}")
print(f"coverage R@5   = {cov5/len(non_abs):.4f}")
print(f"total-miss@50  = {sum(1 for q in non_abs if not (gold[q] & set(rows[q][:50])))}")
EOF

Expected output β€” recall_any@5 = 0.9820 (v0.19.0 file), 0.9840 (v0.17.0) or 0.9820 (v0.16.0 file), total-miss@50 = 0: every question has all its gold sessions somewhere in the top-50; everything below perfect is ranking order, not missing data.

If the file ever stops matching the README, that's a bug β€” open an issue.

2) Run a real subset locally (50 questions)

This actually runs uteke: builds the store from each question's haystack, embeds ~2,400 sessions, and measures what the retriever brings back. Deterministic on the same architecture β€” your numbers should match ours.

cd benchmarks/longmemeval

# 1. Dataset (oracle subset = the 50 hardest-with-known-gold questions, ~5 MB)
./scripts/download_data.sh

# 2. Python deps (stdlib is almost enough; print_metrics.py uses numpy if present)
pip install -r scripts/requirements.txt

# 3. Build uteke from this repo (the harness resolves THIS binary first β€”
#    never a uteke from your PATH)
cd ../.. && cargo build --release -p uteke-cli && cd benchmarks/longmemeval

# 4. Run 50 questions (first run downloads the ~200 MB embedding model once)
python3 scripts/run_eval.py \
    --data data/longmemeval_oracle.json \
    --output results_mine \
    --limit 50

# 5. Score it
python3 scripts/print_metrics.py results_mine/retrieval_results.jsonl

Notes:

  • Runs pinned to 2 cores at low priority β€” your machine stays usable.
  • ~1-2 h for 50 questions on a modern laptop (embedding is the cost).
  • Every question is stored with its own isolated store; no cross-contamination.
  • Want the exact published split instead of oracle? Swap in the full dataset: curl -L -o data/longmemeval_s_cleaned.json \ https://huggingface.co/datasets/xiaowu0162/longmemeval-cleaned/resolve/main/longmemeval_s_cleaned.json then run with --limit 100 (or more) on that file.

3) Full 500-question run (Modal, or any Linux box)

The published numbers come from the fan-out harness on Modal (10 shards Γ— 50 questions, resumed automatically on preemption). Same harness, bigger loop:

cd benchmarks/longmemeval

# dataset (full S split, ~50 MB)
curl -L -o data/longmemeval_s_cleaned.json \
  https://huggingface.co/datasets/xiaowu0162/longmemeval-cleaned/resolve/main/longmemeval_s_cleaned.json

# pin the EXACT binary you want to measure (builds from source inside the image;
# do not skip this β€” an untagged image silently measures whatever it was built with)
export UTEKE_GIT_REF=<commit-sha-or-tag>     # e.g. the v0.19.0 release tag
export MODAL_TOKEN_ID=... MODAL_TOKEN_SECRET=...

modal run scripts/modal_fanout.py --strategy default --num-shards 10

# watch progress (shards appear at completion), then score the merged output:
python3 scripts/print_metrics.py results_v017_revalidation/retrieval_results.jsonl

Total cost last run: under $10 (~2.5 h wall clock, 2 preemptions auto-recovered). No Modal account? The same run_eval.py runs on any Linux box β€” it just takes longer serially (~4 min/question on a 4-core desktop).


Reading the metrics (one paragraph)

LongMemEval-S hides facts across ~115 chat sessions per question and asks where the evidence sessions rank. We report two families because they answer different questions:

  • recall_any@K β€” "did the retriever surface the evidence?" (at least one gold session in top-K). This is the metric competitor benchmarks publish; ours: 98.2% @ 5 (full 500-question set, v0.19.0).
  • recall_all@K (strict) β€” "did it get all of them?" (every gold session in top-K; 65% of the 500 questions have multiple gold sessions). The honest ceiling-capable number: 88.0% @ 5 on the full 500 (88.3% on the 470 non-abstention) β€” mathematical ceiling is 99.4% (3 questions have 6 gold sessions; top-5 physically can't hold all).
  • coverage R@K β€” partial credit (fraction of gold sessions found). 94.3% full-500 (94.6% on the 470 non-abstention).

Per-category breakdown, the 30 abstention questions, contradiction segment, and the v0.16.0 β†’ v0.17.0 β†’ v0.19.0 comparisons: RESULTS.md, docs/benchmarks.md.

Reproducibility guarantees

  • Deterministic: same binary + same dataset + same architecture β‡’ identical rankings. The harness embeds nothing randomly and stores ranked session ids for every question, so any claim can be re-derived from committed artifacts.
  • Pinned provenance: published runs record the exact git ref baked into the binary (UTEKE_GIT_REF) and the raw JSONL is committed β€” no "trust me".
  • Cross-architecture honesty: ARM vs x86 float noise can swap adjacent ranks on tail questions (documented: 107/108 questions identical in our own re-run, 1 adjacent-rank near-tie with the same top-10 set). Aggregates are stable; individual rank order is stable up to Β±1 near ties.