Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
104 commits
Select commit Hold shift + click to select a range
2a394e5
fix(docs): standardize recall latency to ~45ms
ajianaz Aug 11, 2026
e381b17
Merge pull request #999 from codecoradev/fix/docs-recall-latency-30ms…
ajianaz Aug 11, 2026
cc8a69b
chore: harden .gitignore + restructure repo for public safety
ajianaz Aug 13, 2026
c80655a
feat: hybrid as default recall strategy + import fixes (#1005, #1014)…
ajianaz Aug 13, 2026
4a56cac
perf: Phase 2 — skip update check in batch mode, pre-truncate embeddi…
ajianaz Aug 13, 2026
20a0455
feat: Phase 3 — lifecycle/deprecated, memory tools guide, source prov…
ajianaz Aug 13, 2026
6403f01
feat: scene-segmented extraction with priority scoring (#1009) (#1018)
ajianaz Aug 13, 2026
af0e500
chore: bump version to 0.14.0 (#1019)
ajianaz Aug 13, 2026
87a0280
fix: crates.io publish race condition + uteke-mcp metadata (#1021)
ajianaz Aug 14, 2026
5e063e2
fix(ci): switch deploy-website from npm to bun (#1022)
ajianaz Aug 14, 2026
4aea7c4
feat(test): adopt cargo-mutants for mutation testing (#1023)
ajianaz Aug 14, 2026
5441eab
fix(chunker): heading duplication + multibyte infinite loop; 40 mutat…
ajianaz Aug 14, 2026
6a1ec8e
chore: bump version to 0.14.1 (#1025)
ajianaz Aug 14, 2026
b98f579
fix(cli): uteke-cli crates.io publish failure + release verify step (…
ajianaz Aug 14, 2026
6e4dac1
chore: bump version to 0.14.2 (#1031)
ajianaz Aug 14, 2026
2253e07
ci(release): reject stale RELEASE_NOTES.md (#1033)
ajianaz Aug 14, 2026
9d3ad3a
fix: HTTP & MCP recall strategy — default hybrid, loud validation (#1…
ajianaz Aug 15, 2026
76d24b5
chore: bump version to 0.14.3 (#1039)
ajianaz Aug 15, 2026
c2d85de
fix(ci): release notes heredoc executed installer on runner — use tem…
ajianaz Aug 15, 2026
fdc0c68
docs: sync stale version refs, roadmap v0.13–v0.14, merge comparison …
ajianaz Aug 17, 2026
3780a3e
docs(agent): Critical Rule #14 — explicit approval before execution (…
ajianaz Aug 17, 2026
f5029a3
test: use ORT_LIB_NAME in find_ort_in_dir exact-match test (#1054) (#…
ajianaz Aug 17, 2026
de17609
chore: switch ID generation new_v4 → now_v7 (#1058) (#1060)
ajianaz Aug 17, 2026
ec23866
fix: soft-forgotten memories leak into list, search, and doctor count…
ajianaz Aug 17, 2026
8451a80
fix: search_content(None) coerced to default namespace — cross-namesp…
ajianaz Aug 17, 2026
5b8ec19
fix: recall cache hit skips salience/recency boosts — cold/warm score…
ajianaz Aug 17, 2026
021f22f
fix: /export drops namespace attribution + deprecated-row delta undoc…
ajianaz Aug 17, 2026
f6337fc
fix(mcp): uteke_dream destructive defaults — dry-run first, scope or …
ajianaz Aug 17, 2026
0a6331d
feat(mcp): enrich tool outputs agents need — stats tiers/namespaces, …
ajianaz Aug 17, 2026
faaf7f3
feat(mcp): uteke_get + uteke_update tools; short-ID resolution for al…
ajianaz Aug 17, 2026
b8c695e
feat: structural export — full-store round-trip (rooms, graph, edges,…
ajianaz Aug 17, 2026
33fad47
feat: supersession workflow — mark stale decisions superseded, flagge…
ajianaz Aug 17, 2026
db7299a
ci: add source-branch-check — PRs into main must come from develop (#…
ajianaz Aug 19, 2026
5983654
chore: bump version to 0.15.0 (#1073)
ajianaz Aug 19, 2026
072e188
ci: source-branch-check allows chore/release-* release branches (#1075)
ajianaz Aug 19, 2026
2102724
ci: bump actions/download-artifact v4 -> v8 (Node 24) (#1080)
ajianaz Aug 19, 2026
2076bc0
fix(core): structural import silently dropped room_documents + harden…
ajianaz Aug 19, 2026
38a5db8
fix(ci): skip CLA check for bots (dependabot, renovate, github-action…
ajianaz Aug 20, 2026
0dd4592
feat(core): add author_type (human|agent) to memory records (#1084)
ajianaz Aug 24, 2026
0ef1854
fix(server): wire 'at' time-travel param into /room/recall (#1085)
ajianaz Aug 24, 2026
d6ff6bb
docs(server): document deprecation time-travel limitation (#1086) (#1…
ajianaz Aug 24, 2026
bca7666
feat(core): room semantic segmentation (LLM-free, #1088) (#1090)
ajianaz Aug 25, 2026
1a61919
feat(core): expose semantic segments in room_summary (#1088) (#1091)
ajianaz Aug 25, 2026
5052c32
feat(core): segment-level consolidation planner, measure-only (#1088)…
ajianaz Aug 25, 2026
5ab4633
feat(core): provenance trust policy for consolidation (#1089) (#1093)
ajianaz Aug 25, 2026
f6bdb84
feat(core): LLM consolidation executor reusing extraction setup (#108…
ajianaz Aug 25, 2026
0c3f89e
feat: wire consolidation pipeline into core, CLI, and HTTP (#1088) (#…
ajianaz Aug 25, 2026
c1bba22
chore(docs): regenerate api-reference via docgen (#1095 follow-up) (#…
ajianaz Aug 25, 2026
9bc82c4
feat(core): vecq pure-Rust vector backend behind feature flag (#1099)
ajianaz Aug 26, 2026
17e2f91
fix(core): track deprecated_at for correct time-travel recall (#1086)…
ajianaz Aug 26, 2026
13b1d6d
chore: remove embed_cache.db artifact and ignore it (#1101)
ajianaz Aug 26, 2026
e7c127e
feat: per-pair dedup control via POST /consolidate/pair (#1076) (#1102)
ajianaz Aug 26, 2026
f07dfef
fix(ci): keep v prefix in release notes download table + pinned insta…
ajianaz Aug 26, 2026
19a1fe1
feat(server): add --version/-V flag to uteke-serve (#1044) (#1104)
ajianaz Aug 26, 2026
72fd7d9
fix(cli): honor UTEKE_HOME as store path override (#1105) (#1107)
ajianaz Aug 26, 2026
1a081e2
fix: echo author_type in remember responses + CLI --author-type flag …
ajianaz Aug 26, 2026
3f0cf32
fix(core): repair() no longer evicts document chunk vectors from inde…
ajianaz Aug 26, 2026
27a2b4f
fix(build): vecq binaries now buildable via forwarded backend feature…
ajianaz Aug 26, 2026
b4f0b8a
fix(core): backend-aware index filename, doctor label, forget --confi…
ajianaz Aug 26, 2026
668ed6c
fix(core): verify/doctor count doc chunks on DB side, no false MISMAT…
ajianaz Aug 26, 2026
9c1979e
fix(cli): repair report compares against memories + chunks (#1117)
ajianaz Aug 26, 2026
46e1c47
test(longmemeval): fast-eval foundation — datasets + Modal harness + …
ajianaz Aug 27, 2026
e65f36a
test(longmemeval): temporal date-window boost (#1119) (#1126)
ajianaz Aug 27, 2026
ac951f7
test(longmemeval): MMR diversity rerank flag + replay recording (#112…
ajianaz Aug 27, 2026
4e32746
fix(longmemeval): dataset-tagged volume paths + qid-set cache validat…
ajianaz Aug 27, 2026
9cb7cbb
fix(core): make sha2 non-optional — no-default-features profile compi…
ajianaz Aug 28, 2026
e27dbe5
feat(core): Fusion as default recall strategy + benchmark tooling (v0…
ajianaz Aug 28, 2026
c6cf669
feat(bench): UTEKE_GIT_REF — build validation binary from exact SHA i…
ajianaz Aug 29, 2026
bfbc296
fix(bench): install build-essential for git-ref image build (linker c…
ajianaz Aug 29, 2026
14b37ab
docs(bench): pure-default 500Q validation — fusion 0.946 R@5 (+9.2 vs…
ajianaz Aug 29, 2026
8a6d48f
docs: fusion weights = internal hasil tuning, bukan bagian kontrak pu…
ajianaz Aug 29, 2026
8dc6c4a
docs(bench): LongMemEval-S 500Q results page — dual-metric + comparis…
ajianaz Aug 29, 2026
bfed150
chore(release): CHANGELOG 0.16.0 — full-release validation numbers + …
ajianaz Aug 29, 2026
468f842
fix(core): repair --reembed can now repair NULL-embedding rows (#1146…
ajianaz Aug 29, 2026
7facdeb
fix(ci): CLA check skips bot PRs with exact allowlist (#1132) (#1148)
ajianaz Aug 29, 2026
18ff623
fix(cli): repair --reembed report reflects final index state (#1149) …
ajianaz Aug 29, 2026
da46d40
chore(ci): bump actions/github-script from v7 to v9 (#1151)
ajianaz Aug 29, 2026
febee27
test(cli): resolve_namespace precedence — flag > env > config > defau…
ajianaz Aug 29, 2026
512f379
test(cli,core): env parsing guards + validate_input_with_limits bound…
ajianaz Aug 29, 2026
d03375c
test(core): ensure_embedder branch coverage — unknown/custom/openai-m…
ajianaz Aug 29, 2026
978003f
test(cli): cover TOML helpers — migrate_content, global_config_path, …
ajianaz Aug 29, 2026
a0a400d
docs: refresh documentation for 0.16.0 — fusion default, updated benc…
ajianaz Aug 29, 2026
92b485b
chore: sync main back into develop after #1158 docs carry (#1159)
ajianaz Aug 29, 2026
a83987d
docs(contributing): require CLA + state contributions are unpaid (#1162)
ajianaz Aug 30, 2026
cf49be8
deps(core): bump vecq-core 0.1.1 -> 0.3.0, adopt 4-bit + residual pro…
ajianaz Aug 31, 2026
10ddcc4
fix(core): feature-aware embedder default in open() + open_with_backe…
ajianaz Aug 31, 2026
75b72ff
feat(core): runtime-selectable vector engine (usearch + vecq in one b…
ajianaz Aug 31, 2026
c0b92cb
feat(build): dual-engine default — official binaries ship usearch + v…
ajianaz Aug 31, 2026
652fca6
docs(bench): independent reproduction results + surfaced headline num…
ajianaz Sep 1, 2026
4dd7c54
chore(deps): bump flate2 from 1.1.9 to 1.1.10 (#1175)
dependabot[bot] Sep 3, 2026
39f877a
chore(deps): bump which from 8.0.5 to 8.0.6 (#1176)
dependabot[bot] Sep 3, 2026
2cc0474
chore(deps): bump usearch from 2.26.0 to 2.26.1 (#1177)
dependabot[bot] Sep 3, 2026
51707f3
chore(deps): bump uuid from 1.24.0 to 1.26.0 (#1178)
dependabot[bot] Sep 3, 2026
bd2288c
fix(room): honest room_delete reporting across CLI/HTTP/MCP — count u…
ajianaz Sep 3, 2026
cbccebb
fix(server): resolve memory IDs to graph nodes in /graph/edge endpoin…
ajianaz Sep 5, 2026
0f75a96
feat(api): namespace management — move, rename/merge, delete with str…
ajianaz Sep 5, 2026
16f1186
feat(core): provenance data model — source hash, event actor/evidence…
ajianaz Sep 5, 2026
841e990
feat: contradiction resolution ledger + undo across CLI/HTTP/MCP (#11…
ajianaz Sep 5, 2026
d0243a7
feat: explain recall — ranking signals on CLI, HTTP, and MCP (#1160) …
ajianaz Sep 6, 2026
8a28aef
fix: graph excludes stale memory nodes + readable node labels (#1189 …
ajianaz Sep 6, 2026
9bd004d
feat: /list pagination metadata envelope (#1188) (#1191)
ajianaz Sep 6, 2026
29d50d7
feat: contradiction benchmark segment + uteke supersede CLI (#1172 F3…
ajianaz Sep 6, 2026
b15c2c4
chore: release 0.17.0 (#1194)
ajianaz Sep 6, 2026
fef28c7
merge: main into develop (0.17.0 release prep)
ajianaz Sep 6, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 4 additions & 0 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -110,6 +110,10 @@ jobs:
# #1131: the embedder-less profile (cloud/OpenAI embedders only, no
# onnx) must keep compiling — the embedding cache is backend-agnostic.
- run: cargo check -p uteke-core --no-default-features --features vecq
# #1168: dual-engine profile (usearch + vecq compiled together, runtime
# selection) must keep compiling and passing the engine-switch tests.
- run: cargo test -p uteke-core --features "usearch,vecq" --lib vector_engine
- run: cargo test -p uteke-core --features "usearch,vecq" --lib memory::vector

docs-check:
name: API Docs Fresh
Expand Down
6 changes: 3 additions & 3 deletions .github/workflows/cla-check.yml
Original file line number Diff line number Diff line change
Expand Up @@ -12,7 +12,7 @@ permissions:
jobs:
cla-check:
runs-on: ubuntu-latest
if: "!contains(fromJSON('[\"app/dependabot\", \"app/renovate\", \"github-actions[bot]\"]'), github.event.pull_request.user.login)"
if: "!contains(fromJSON('[\"dependabot[bot]\", \"renovate[bot]\", \"github-actions[bot]\", \"app/dependabot\", \"app/renovate\"]'), github.event.pull_request.user.login)"
steps:
- name: Fetch & check CLA signature
id: check
Expand Down Expand Up @@ -48,7 +48,7 @@ jobs:

- name: Comment on PR (unsigned only)
if: steps.check.outputs.signed != 'true'
uses: actions/github-script@v7
uses: actions/github-script@v9
with:
script: |
const author = '${{ github.event.pull_request.user.login }}';
Expand Down Expand Up @@ -97,7 +97,7 @@ jobs:
}

- name: Set commit status
uses: actions/github-script@v7
uses: actions/github-script@v9
with:
script: |
const signed = '${{ steps.check.outputs.signed }}' === 'true';
Expand Down
29 changes: 29 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,34 @@
# Changelog

## [0.17.0] — 2026-09-06

Minor release. Theme: **inspectable, trustworthy memory** — explain recall on
every surface, auditable conflict resolution with a measurable payoff, honest
graphs, pagination metadata, and a dual-engine vector layer.

### Added

- **Contradiction benchmark segment (#1172, phase 3)** — `benchmarks/longmemeval/contradiction_segment.py`: 40-topic active-store segment measuring conflict-resolution quality end-to-end. Baseline (both facts active) vs resolved (superseded): fusion winner@1 0.975 → 1.000, stale@5 0.825 → 0.000. Published in `benchmarks/longmemeval/RESULTS.md`. Also adds `uteke supersede <old> <new> [--reason]` — CLI surface parity for supersession (previously MCP/HTTP only).

- **`/list` pagination metadata (#1188)** — `POST /list` accepts `"include_meta": true` to respond with an envelope `{memories, total, has_more, next_offset}` (`next_offset` is `null` on the last page) so clients no longer blind-paginate with 100-row guesses. The default response is unchanged (bare array) — existing clients are untouched; `include_meta` is ignored in `at` (point-in-time) mode, which stays a bare array.

- **Explain recall (#1160)** — `explain` mode on every recall surface shows WHY each memory ranked where it did: vector similarity and rank, FTS rank, RRF score with per-channel fusion contributions, and jaccard/salience/recency/graph boost deltas. Surfaces: `uteke recall "…" --explain` (human-readable, combine with `--json` for machine output), `POST /recall` with `"explain": true` (memory-only — combined with `search_type`/`at`/`before`/`after` returns 400), and the `explain` flag on the MCP `uteke_recall` tool. The explanation path replays the active strategy's exact pipeline (same channel depths, RRF constants, and boost order) while bypassing the recall cache, so the explanation always matches the returned results; fts5 explanation works without an embedder, other strategies embed the query once (~50 ms, same as a cold recall).

- **Contradiction resolution ledger + undo (#1172, phase 2)** — supersessions are now a first-class, auditable ledger instead of a side effect: `Uteke::contradiction_resolutions(namespace, limit)` lists superseded-but-not-restored memories (winner, reason, timestamp via the deprecation row), `Uteke::undo_supersession(id)` restores a retired memory, removes the supersession edge pair, and records a `supersession_undone` event on both sides (only memories carrying a live `superseded_by` edge can be undone — the undo is itself auditable). Ledger membership is edge-driven (deprecated row + `superseded_by` edge), the same predicate undo resolves against, and re-superseding an already-deprecated memory refreshes the stored reason/timestamp so the ledger always names the current winner. Surfaces: `GET /contradictions?namespace=&limit=`, `POST /contradictions/undo` (`{id}`; 404 when nothing to undo), `uteke contradictions list|undo`, and MCP `uteke_contradictions` / `uteke_contradictions_undo`. Fixed in the process: the no-namespace ledger query bound its limit parameter to a nonexistent placeholder (`?2`) and failed at runtime — caught by the new MCP roundtrip test.

- **Provenance data model (#1172, phase 1)** — schema v18 (additive): `memories.source_hash` records the SHA-256 of content at write time (tamper evidence — audits recompute it against live content), and `timeline_events.actor`/`evidence_json` record who performed an event and what evidence supports it. New `Uteke::provenance(id)` returns the full report (provenance fields, trust tier, hash comparison, event chain) — exposed as `GET /provenance?id=`, `uteke provenance <id>`, and the `uteke_provenance` MCP tool.

### Fixed

- **Graph data returned stale nodes (#1189)** — `GET /graph` without a namespace returned every `graph_nodes` row raw, including nodes whose parent memory had been forgotten or deprecated; with soft-delete the store accumulated stale nodes on every conflict resolution. Memory-linked nodes are now filtered by liveness (memory exists and `deprecated = 0`) in every `graph_data` path, edges touching removed nodes are dropped, and `stats` counts the filtered graph.
- **Memory graph nodes labeled with raw UUIDs (#1187)** — `ensure_node_for_memory` now labels new memory nodes with a readable content preview (first 60 chars of the memory) instead of the raw memory UUID, and upgrades legacy UUID-labeled rows in place on next access. Entity nodes are unaffected.

- **Namespace management API (#1181)** — namespaces are a derived view, now with sanctioned ops: `PUT /memory` accepts `namespace` (move a memory — plain column update, no re-embed), `POST /namespaces/rename` (`{from, to}`; existing target = merge, returns `{from, to, moved, target_existed}`), and `POST /namespaces/delete` with an explicit strategy for its memories: `refuse` (default — 409 while any memory references the name), `merge` (move all memories to `target`, the name vanishes), or `deprecate` (soft-delete — restorable via promote, never hard-deleted). `GET /namespaces?with_counts=true` now adds `active`/`deprecated` breakdown fields (`count` stays the total). CLI parity: `uteke namespace move|rename|delete` (delete requires `--confirm`). MCP parity: `uteke_namespace_rename`, `uteke_namespace_delete`, and `namespace` field on `uteke_update`.

### Fixed

- **`POST /graph/edge` always returned 500 for valid memory IDs (#1180)** — the handler validated `source`/`target` as memory IDs but inserted them directly into `graph_edges`, whose foreign keys point at `graph_nodes(id)`. Memory IDs are now resolved to their linked graph node (or a node is ensured automatically) before insertion. `DELETE /graph/edge` accepts memory IDs or graph node IDs the same way, and its documented query params are corrected to `?source=...&target=...`. `POST /graph/edge` now responds with `{ok, source_node, target_node}` so clients can track the created nodes.

## [0.16.0] — 2026-08-28

Minor release. One theme: retrieval quality that ships by default.
Expand Down
26 changes: 25 additions & 1 deletion CONTRIBUTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -272,4 +272,28 @@ See [CODE_OF_CONDUCT.md](CODE_OF_CONDUCT.md).

## License

By contributing you agree your work is licensed under [Apache-2.0](LICENSE). No CLA required.
By contributing you agree your work is licensed under [Apache-2.0](LICENSE).

## CLA

All contributions (code, docs, tests, configuration) require a signed
Contributor License Agreement before a pull request can be merged:

- 📋 **Individual?** → [Sign the Individual CLA](https://codecoradev.github.io/cla/?type=individual)
- 🏢 **Contributing on behalf of a company?** → [Sign the Corporate CLA](https://codecoradev.github.io/cla/?type=corporate)

The CLA is a license agreement, not a copyright assignment — you keep
ownership of your work. Signing takes a couple of minutes and is stored
in the [codecoradev/.github](https://github.com/codecoradev/.github)
repository; a bot checks it automatically on every pull request.

## Contributions are unpaid

Contributing to this project is **voluntary and unpaid**. There is no
compensation, payment, bounty, or financial reward of any kind for
contributions — now or in the future. You contribute on your own time,
at your own discretion, because you want to improve the project.

If any paid-contribution program is ever introduced, it will be announced
explicitly and this document will be updated. Until then, assume every
contribution is volunteer work under the Apache-2.0 license terms above.
39 changes: 23 additions & 16 deletions Cargo.lock

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

2 changes: 1 addition & 1 deletion Cargo.toml
Original file line number Diff line number Diff line change
Expand Up @@ -9,7 +9,7 @@ members = [
]

[workspace.package]
version = "0.16.0"
version = "0.17.0"
edition = "2024"
license = "Apache-2.0"
repository = "https://github.com/codecoradev/uteke"
Expand Down
8 changes: 7 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -142,7 +142,13 @@ Every AI tool forgets. Context windows fill up, sessions end, and your AI starts
| **Recall latency (10K memories)** | **42ms** P50, 50ms P95 | Flat from 100 to 10K memories (HNSW O(log N)) |
| **Insert throughput** | 6-22 ops/s | CPU-bound (ONNX embedding inference) |
| **Storage per memory** | ~10KB | SQLite + HNSW, scales linearly |
| **LongMemEval Recall@5** | **0.946** | Full 500Q validation, zero-config fusion default (R@10 0.977), EmbeddingGemma Q4 |
| **LongMemEval-S recall_any@5** | **98.2%** | Full 500Q validation, zero-config fusion default (v0.16.0) — the metric competitor benchmarks publish |
| LongMemEval-S recall_all@10 | 95.4% | Strict: every gold session in top-10 |
| LongMemEval-S strict recall_all@5 | 88.0% | Every gold session in top-5 (mathematical ceiling 99.4%) |

![uteke vs published systems on LongMemEval-S](docs/assets/longmemeval-comparison.jpg)

> **Don't trust our benchmark — run your own.** We re-ran 108 of the 500 published questions on a 4-core ARM desktop (different CPU architecture from the published Modal x86 run, same v0.16.0 binary and harness): **107/108 produced identical per-question rankings**. The single difference was an adjacent-rank near-tie, both runs retrieved the identical top-10 session set, one gold session swapped ranks 5-6. Details in [RESULTS.md](benchmarks/longmemeval/RESULTS.md).

Full benchmarks: `uteke bench --counts 100,1000,10000 --json` · [Benchmark details](docs/benchmarks.md) · [LongMemEval results](benchmarks/longmemeval/RESULTS.md) — fusion default: R@5 0.946 / R@10 0.977 on the full 500Q validation set (v0.16.0)

Expand Down
64 changes: 64 additions & 0 deletions benchmarks/longmemeval/RESULTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,28 @@

---

## Independent Reproduction (2026-09-01)

The published 500Q run was produced on Modal x86. To test whether the result depends on that infrastructure, a 108-question subset (pref: 30, kupd: 78) was re-run locally on a 4-core ARM desktop (Oracle Ampere A1, aarch64) — same v0.16.0 binary, same harness (`run_eval.py`), different CPU architecture.

| Subset | Questions | Identical per-question rankings | R@5 published | R@5 re-run |
|---|---|---|---|---|
| pref | 30 | 30/30 | 96.7% | 96.7% |
| kupd | 78 | 77/78 | 100.0% | 99.4% |
| **Total** | **108** | **107/108** | — | — |

**The one divergence** (question `0977f2af`, knowledge-update, 2 gold sessions): both runs retrieved the identical top-10 session set; one gold session sits at rank 5 (published) vs rank 6 (re-run), moving that question's recall_all@5 from 1.0 to 0.5. The harness persists rankings rather than raw scores, so the exact score gap cannot be shown, but the identical top-10 set identifies this as cross-architecture floating-point noise on an RRF near-tie, not a retrieval failure. R@10 = 1.0 in both runs.

**Reproduce:**

```bash
python3 run_eval.py --data data/subset_kupd.json --output results_rerun --strategy default --resume
```

Raw artifacts are kept on the benchmark Modal volume (`uteke-longmemeval`, `default/` and rerun prefixes), consistent with the published run — datasets and result JSONL files are not committed to git (see `.gitignore` here; `download_data.sh` fetches the dataset).

---

## Uteke Retrieval — Strategy Comparison

### Vector (semantic only)
Expand Down Expand Up @@ -115,3 +137,45 @@ Binary built in-image from exact SHA bfbc296 (PR #1137/#1138), image build print

Context: v0.15.0 hybrid baseline on the same dataset: Overall R@5 = 0.854 / R@10 = 0.885 (2026-08-13).
Fusion default lifts full-500Q recall@5 by **+9.2 points** (0.854 → 0.946) with zero configuration.

---

## Contradiction-resolution segment (#1172 Fase 3) — 2026-09-06

**Active-store knowledge-update segment**: 40 topics × (stale fact + winner fact + 3 distractors),
queries ask "which {thing} does {topic} use now?" (semantic, no keyword echo of the answer).
Baseline ranks with BOTH facts active; resolved ranks after `supersede(stale → winner)` —
baseline is measured for every strategy BEFORE any resolution, then the store is resolved once.
Binary: local release build (0.16.0 + #1185 ledger), local ONNX EmbeddingGemma, ARM64.

Harness: `contradiction_segment.py` (this directory). Raw metrics: `results_contradiction_f3/metrics.json`.

| Strategy | Stage | winner@1 | winner@5 | winner MRR | stale@1 | stale@5 |
|---|---|---|---|---|---|---|
| fusion (default) | baseline (unresolved) | 0.850 | 1.000 | 0.925 | 0.150 | **1.000** |
| fusion (default) | resolved (superseded) | **1.000** | 1.000 | **1.000** | 0.000 | 0.000 |
| hybrid | baseline | 0.025 | 1.000 | 0.469 | 0.975 | 1.000 |
| hybrid | resolved | 0.225 | 1.000 | 0.588 | 0.000 | 0.000 |
| vector | baseline | 0.950 | 1.000 | 0.975 | 0.050 | 1.000 |
| vector | resolved | 1.000 | 1.000 | 1.000 | 0.000 | 0.000 |

Findings:

- **Unresolved conflicts pollute every strategy's top-5**: with both facts active, the stale
fact sat in top-5 for 100% of topics on all strategies (hybrid's BM25 even ranks the stale
fact top-1 for 97.5% of topics — the old fact's "uses X" phrasing matches "use now?"
queries lexically). After `supersede`, stale@1 and stale@5 drop to **0.000** everywhere
(deprecated memories are excluded from recall).
- **Supersede lifts the default surface**: fusion winner@1 0.850 → 1.000, MRR 0.925 → 1.000;
vector 0.950 → 1.000. Hybrid stays weakest on winner@1 (lexical BM25 keeps the new fact's
"switched to" phrasing behind distractors) but its stale pollution is fully cleared.
- **Ledger integrity**: `contradictions list` listed all 40 resolutions; every stale fact is
restorable via `contradictions undo` (auditable conflict resolution, #1172 F2).

Interpretation: ranking alone often picks the winner, but only explicit conflict resolution
guarantees stale facts leave the retrieval surface — the difference between "usually right"
(85–95% top-1) and deterministic freshness (100% top-1, zero stale). For agent memory, where
"use now" queries are the norm, resolution is what keeps top-1 trustworthy. This segment is
synthetic and deterministic (fixed topic list); it measures the conflict-resolution pipeline,
not LongMemEval dataset recall.

Loading
Loading