Skip to content

feat(storage): operator command to inspect and apply offline store upgrades - #56

Merged
Knucklessg1 merged 3 commits into
mainfrom
feat/operator-store-upgrade
Oct 6, 2026
Merged

Knucklessg1 merged 3 commits into
mainfrom
feat/operator-store-upgrade

Conversation

@Knucklessg1

Copy link
Copy Markdown
Member

What

The storage kernel has offline, data-preserving owner-store upgrades, but nothing except tests could run them, and the refusal of an upgradable store told the operator to move the file aside. This adds the entry point:

  • one registry of offline upgrades in the storage kernel;
  • an operator command on the server binary, store-upgrade inspect <data-dir> and store-upgrade apply <data-dir> --confirm;
  • a startup refusal that names the exact command. The engine never upgrades a store itself.

No existing upgrade's semantics, digests, lineage declarations or on-disk formats change, and nothing is upgraded at an ordinary open.

Wire contract: unchanged. The refusal stays what it already was, a named {LAYOUT}_FORMAT_UPGRADE_REQUIRED error declared by the lineage registry and reported in the startup log with exit status 1; no generated contract artifact moves.

Inventory (what was on main)

Store kind (file) Predecessors with an offline upgrade inspect → apply
SQL catalog (sql-catalog/*.redb) before durable source checkpoints; before durable ANN and edge index generations inspect_sql_source_checkpoint_upgrade (and its two aliases) → upgrade_sql_source_checkpoints
Agent Library (agent_library.redb) before served MCP catalog authority inspect_agent_library_mcp_catalog_upgrade → upgrade_agent_library_mcp_catalog
Graph shard (graph-<n>.redb) before repository enrichment; before the audit-append idempotency index inspect_graph_shard_upgrade(…, source) → upgrade_graph_shard

Declared predecessors with no upgrade (still refused with the move-aside step): Agent Library before connector packs, blob store before holder-scoped references, graph shard before the storage scrub, graph shard before enrichment policy revisions.

Token and confirmation model. Every inspection takes the file path, the physical identity the caller expects, an optional private-payload authenticator, and a private local staging root with a byte budget. It pins the source file, copies it to an anonymous scratch file, lets native recovery run on the copy, and checks the pinned predecessor digest, the exact table census and the typed row evidence. It returns a token that cannot be cloned or constructed elsewhere, bound to the file descriptor and a SHA-256 of the bytes. Applying the token takes the file exclusively (and proves the filesystem enforces the lock), re-checks all of that, and performs one immediate-durability commit that creates the missing empty tables and replaces the owner manifest (authority epoch + 1). Existing rows are not rewritten.

What they refuse: any other generation (the current one included), a different physical identity, a stray or multimap table, bytes that changed after inspection, a staging root under /tmp or .cache or one that is not private, and non-Linux hosts.

No backup. The upgrade functions keep no copy of the source: the "quarantine" in that code retires the scratch file, not the store. The safety net is the single atomic commit (main already has a crash-stage test for it).

How the server reported such a store. Every manifest read refuses a declared predecessor with {LAYOUT}_FORMAT_UPGRADE_REQUIRED and tells the operator to move the file aside, upgradable or not; any other digest is OWNER_STORE_FORMAT_UNKNOWN. Graph shards fail at boot. The Agent Library and the SQL catalogs open lazily, and the SQL catalog path replaces the reason with "could not be opened".

Operator CLI surface. The server binary had flags only. The offline tools are separate binaries shipped beside it (migrate-shards, restore): one clap parser each, the persistence directory from --persist-dir or GRAPH_SERVICE_PERSIST_DIR, the engine stopped, one result line, exit status 1 on failure. Each of them links the whole engine, so a third would add another engine-sized binary to the published package, and the package payload has its own owner. The command is therefore a subcommand of the server binary, following the same conventions: no packaging change, and the tool is by construction the same build as the engine that will open the stores afterwards.

Design

Registry — crates/eg-storage/src/owner/offline_upgrade.rs. OFFLINE_STORE_UPGRADES is a table of rows: declared predecessor (which names the store kind, the generation and the file), inspect function, apply function. From a row follow the read-only classification (classify_owner_store_format: current / upgrade available / declared predecessor with no upgrade / unknown digest), the refusal text, and the operator document.

Command — src/server/persistence/store_upgrade.rs, wired as a clap subcommand in src/operator_command.rs.

  • inspect <data-dir> reads each store's manifest read-only and writes nothing (not even the lock file).
  • apply <data-dir> --confirm takes the engine's own single-writer lock (engine.lock, the lock the engine holds for its lifetime) for the whole run, so it refuses while an engine is running and no engine can start meanwhile. For each store with a registered upgrade it runs the kernel's inspect and then its single commit. It stops at the first failure; it never writes a store it does not upgrade; a second run reports nothing to do. No interactive prompt.
  • Stores are found through the existing registries: graph shards by index, the durable-store registry by name, SQL catalogs by directory. The expected physical identity comes from the same declarations the server opens each store with.
  • A store left by an unclean shutdown cannot be read read-only. inspect reports it as not inspectable; apply lets the registered upgrades of that store kind inspect it (they recover a private copy, never the file) and upgrades it if one admits it.
Exit status Meaning
0 nothing to do, or every applicable upgrade was applied
10 (inspect) an upgrade is available
20 a store is in a format this build neither opens nor upgrades; it was not modified
30 a store could not be read and no upgrade admitted it; it was not modified
1 refused (engine running, no --confirm, bad directory) or an upgrade failed and the run stopped

The last output line is one JSON object (command, verb, outcome, exit_code, pre_upgrade_copy, not_processed, error, stores[]).

Startup refusal. Before the first store is opened, the server classifies the data directory with the same function and refuses, with exit status 1, if a registered upgrade applies to any store. This also covers the lazily opened stores. The message is the store's named error followed by Run: epistemic-graph-server store-upgrade apply <the directory> --confirm. The kernel's own refusal of an upgradable predecessor now names the command too, instead of advising to move the file aside.

What a new upgrade adds

One row in OFFLINE_STORE_UPGRADES, plus a predecessor fixture for the command test. The graph-shard upgrade that merged while this branch was open (before the audit-append idempotency index) is registered here exactly that way:

offline_upgrade!(
    GRAPH_SHARD_BEFORE_AUDIT_REQUESTS,
    |path, identity, integrity, options| inspect_graph_shard_upgrade(
        path, identity, integrity, options, GraphShardSource::BeforeAuditRequests
    ),
    upgrade_graph_shard
),

The identity operation family (#46) is a different case, and one row is not enough for it yet. It keeps its new state inside the existing access-control policy image, so the store's table set and layout digest do not change, and nothing can tell an earlier file from a current one. The lineage identifies a predecessor by its owner-table set. To get a fail-closed path through this command, #46 must add:

  1. a change to the access-control store's table set (for example a table for the identity state), with the earlier set declared in the lineage as a predecessor and its digest pinned, so an earlier file is refused by name instead of being read as current;
  2. an inspect/apply pair for that predecessor in the storage kernel: one more target in the existing add-the-missing-tables transition if the new table starts empty, or new code if the stored policy image itself must be rewritten;
  3. one registry row, and a predecessor fixture for the command test.

The command already knows the physical identity of the access-control store (every_registered_upgrade_is_one_the_command_can_run fails if a registered store kind has none), so nothing in the command itself needs to change.

Checks

Rust was built and tested on the build host from a local-disk copy of the branch, because the offline-upgrade tests cannot run from a network mount (the scratch-file reservation is refused there).

Command Result
cargo fmt --all -- --check clean
cargo clippy -p eg-storage --all-targets -- -D warnings clean
cargo clippy -p epistemic-graph --no-default-features --features full,ast-extended --all-targets -- -D warnings clean
cargo clippy -p epistemic-graph --no-default-features --features server --all-targets -- -D warnings clean
cargo test -p eg-storage test result: ok. 120 passed; 0 failed (plus 5 doc tests)
cargo test -p epistemic-graph --no-default-features --features full,ast-extended --lib -- store_upgrade persist_lock durable_stores sql_tables redb_layout agent_library:: test result: ok. 36 passed; 0 failed
cargo test -p epistemic-graph --no-default-features --features full,ast-extended --bin epistemic-graph-server test result: ok. 6 passed; 0 failed
cargo run -p eg-storage --example gen_owner_store_formats document regenerated; the freshness test in the storage suite passes
built binary: store-upgrade inspect <empty dir>, apply without and with --confirm, a missing directory, --help exit 0 / 1 / 0 / 1 as designed; JSON summary on the last line

Hooks, run over origin/main..HEAD: the commit-stage set (33 passed, the rest have no files in scope), and the manual-stage complexity-staged, kiss-changed-rust, dupehound-changed-functions, jscpd-differential, kiss-census, cccc-census, rust-arch-lint, orphan-modules, durable-table-registration and registry-test-ownership: all pass. scripts/check_public_specs.py and scripts/security/check_secret_history.py --base origin/main pass.

Tests added (disposable stores only):

  • storage kernel (owner::offline_upgrade::tests): every registry row is a declared predecessor and is registered once; the refusal names the command only where an upgrade is registered and writes nothing; classification names current / upgradable / no-upgrade / unknown without writing; every row upgrades a genuine predecessor to the current layout exactly once.
  • command (server::persistence::store_upgrade::tests): inspect on current stores reports nothing to do and creates nothing; inspect on a predecessor reports the upgrade and leaves the bytes alone; apply upgrades two graph-shard generations, the Agent Library and a SQL catalog with the exact row bytes kept, and a second apply changes nothing; both verbs refuse while the engine lock is held; apply without --confirm opens nothing; an unknown format and a predecessor with no upgrade are reported (exit 20) and stay byte-identical; a refused upgrade stops the run and the later store is untouched; a predecessor left unclean is not inspectable read-only and is still upgraded by apply; the startup check, the Agent Library open and the graph store open refuse a predecessor with the named error and the command.
  • lock probe (persist_lock::tests) and command-line parsing (operator_command::tests).

Not run here: the whole epistemic-graph test suite (left to CI), and the contract generator check, because no contract input is touched. --features server,security without query does not pass clippy on main today (two unused imports in files this change does not touch), so that combination was not usable as a third profile.

Limits

  • No pre-upgrade copy. The command does not add one: a store file is bound to its inode, so a byte copy could not be put back without a separate rebind step. The output says "pre_upgrade_copy":"none" and the procedure tells the operator to take a filesystem snapshot first if one is wanted.
  • Startup now refuses on an upgradable lazily opened store. That is the requested behavior, and it is a deliberate exception to "a lazily opened store's failure stays local" (EG-DURABLE-KERNEL-R006): only a store that a registered upgrade applies to refuses startup this way.
  • A SQL catalog that is both a predecessor and left unclean is not caught by the startup check (its manifest cannot be read read-only); apply upgrades it, and the lazy open still hides the reason. That reporting path is unchanged here.
  • apply needs free space of about one store's size for the scratch copy, in a private (mode 0700) local directory. The default is created inside the data directory and left empty.
  • After an unclean shutdown every store is unreadable read-only. inspect then exits 30, and apply copies each such store of a kind that has registered upgrades to the scratch directory to find out whether an upgrade admits it. Stores it cannot place are left untouched and reported with exit 30; the engine recovers them at its next start. The deployment job decides whether 30 lets the engine start.
  • Only the upgrades in the registry are covered. The tenant semantic-index migration keeps its own separate path. Stores outside the data directory's top level and sql-catalog/ are not scanned.
  • EG-DURABLE-KERNEL-R070 is recorded as BUILDING with this branch as evidence.

🤖 Generated with Claude Code

Knucklessg1 and others added 2 commits October 2, 2026 16:26
…grades

The storage kernel's offline, data-preserving owner-store upgrades had no
entry point: nothing but tests called them, and the refusal of an upgradable
predecessor told the operator to move the file aside.

- eg-storage: one registry of offline upgrades (declared predecessor,
  inspect function, apply function) and a read-only classification of an
  owner file's format. The named refusal of a registered predecessor now
  names the operator command; a predecessor with no upgrade keeps the
  move-aside step.
- server binary: `store-upgrade inspect <data-dir>` (read-only) and
  `store-upgrade apply <data-dir> --confirm`. Apply holds the engine's own
  directory lock, runs each applicable upgrade through the kernel's
  inspected single-commit transition, stops at the first failure, never
  writes a store it does not upgrade, and is a no-op on a second run. The
  last output line is one JSON object; the exit status distinguishes
  nothing to do, upgrade available, blocked, undetermined and failed.
- startup: the engine refuses a data directory holding a store a registered
  upgrade applies to, with that store's named error and the exact command.
  It never upgrades a store itself.
- the generated owner-store format document carries the operator procedure.

The startup health check for the SPARQL federation endpoint moves from
main.rs to server_startup.rs unchanged, so the subcommand hook does not grow
a file that is already over the size cap.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
EG-DURABLE-KERNEL-R070 is implemented on this branch and not yet merged, so
its delivery state is BUILDING with the implementation commit as evidence.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@Knucklessg1
Knucklessg1 merged commit 8dbb5bd into main Oct 6, 2026
22 checks passed
@Knucklessg1
Knucklessg1 deleted the feat/operator-store-upgrade branch October 6, 2026 03:56
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant