Skip to content

PP.5: five-concurrent BuildingBlock lifecycle, PostgreSQL hardening, fail-closed evidence gates (#1033 #1035 #1036 #1037 #1041 #1043 #1044) - #1046

Open
zoorpha wants to merge 102 commits into
mainfrom
pp5-s5-foundation
Open

zoorpha wants to merge 102 commits into
mainfrom
pp5-s5-foundation

Conversation

@zoorpha

@zoorpha zoorpha commented Sep 23, 2026

Copy link
Copy Markdown
Collaborator

Intent

Build the reviewed PP.5 commit stack on the pp5-s5-foundation branch: 8 commits covering the five-concurrent BuildingBlock lifecycle campaign, the PostgreSQL provider/isolation hardening, fail-closed campaign evidence gates, and the deterministic-timing sweep.

Scope

Files: 45 changed, ~6800 insertions.

Product / deployment profile

  • Profile: small-edge-cloud (minimum target hypervisors corrected to 1).
  • Backends: SQLite (portable default) + PostgreSQL (production) — both proven on the pg lanes.
  • Authority: O3K-owned compute/network/storage resources; execution provider is the libvirt-adjacent bounded provider (compute provider is a fake in the campaign harness).

Non-goals (explicit)

  • No S1/S2/S3/S10/S20 declarations; no soak claims.
  • No HA / SLA claims (no durable work-ownership fencing beyond the existing evidence).
  • No migration/evacuation features.
  • No release publication or version bump.
  • No Araf external-consumer provisioning.
  • No PP.7.

OpenStack compatibility impact

  • Preserves operation-level OpenStack compatibility at the /v2.1 / /v2.0 contract edges in the pp4/pp5 harness; no new advertised OpenStack service surface. O3K internal architecture (shared Cloud Kernel) is unchanged.

Database / footprint

  • No schema migration added on top of existing; terminalization and repair operate on the existing operations / resources rows. PostgreSQL PG + SQLite lanes both green.

Authority model / workflow / compensation

  • Orphan repair and every delete release seat serialized on orphan_repair_lock (leaf-first ordering: orphan_repair_lock → network_mutation_lock → NetworkService lock → store). Replayed/late terminal updates and delete replays are fenced against live re-attachment; deterministic o3k:orphan-unbind op id with fail-closed retry.

Evidence / tests

Run locally on e03223df (RTK-disabled, real PostgreSQL via O3K_DATABASE_URL):

  • cargo fmt --all -- --check — pass
  • cargo clippy --workspace --all-targets --all-features -- -D warnings — pass
  • cargo test --workspace --all-features --no-fail-fast (with PG) — 1748 passed / 0 failed
  • pp4 endpoint-lifecycle lanes — 18 SQLite + 23 PG (include-ignored)
  • PP.5 evidence validator — 27 PASS (pp5_evidence_validation.sh)
  • Guard suites — all green (scale-composition, real-host-workflow, keycloak-authority, dispatch, protected-preflight, postgres-cleanup, libvirt-pool)
  • New [PP.5][Compute/Network] Reconcile server-owned endpoints orphaned by an interrupted terminal delete #1035 regressions (replayed terminal delete preserves a live re-attached port; sweep dispatches at most one unbind per pass) — pass on both lanes

Adversarial review

The endpoint-repair commits received adversarial review against TOCTOU / stale-binding / cross-project-identity / re-attachment / sweep-capacity attack vectors. Two findings were fixed and re-reviewed clean over commits fce0adc5 → 9b3ff63d (create-vs-sweep serialization) and → e03223df (fence every release seat; sweep cap; atomic binding fence). (Note from executor: one independent reviewer could not be spawned in my subagent environment — the review was performed with the same structured methodology and is recorded in the commit bodies.)

Request

Review requested; attention commits: fce0adc5, 9b3ff63d, e03223df.

Lifecycle success previously committed the operation terminal row and the
resource terminal projection as two sequential durable writes; a crash
between them left exactly one half committed, which no reconciliation pass
repairs. terminalize_lifecycle now commits the operation REVERSIBLE->FINAL
transition, the canonical metadata row, and the fenced (generation=expected)
terminal resource projection in ONE transaction on both adapters (SQLite
BEGIN IMMEDIATE; PG pool.begin + FOR UPDATE), with an equivalent replay as a
no-op and a conflicting terminal -> Corrupt.

This commit also makes the tombstone-visibility contract explicit (:89) and
fixes the PostgreSQL adapter that violated it: list_resources_by_kind returns
every resource of the kind including DELETED tombstones on both adapters, so
repair scans (which exist precisely to reconcile tombstones) see durable
terminal truth. Endpoint/network release stays outside the transaction,
repaired by the orphan-endpoint sweep (#1035).
Ownership and repair authority for o3k-server endpoints. The request-path
releases (native + compat delete handlers) and the new repair-only sweep
now re-read the live-attached endpoint set immediately before releasing and
refuse to hand back any port the compute layer still references (a NEW live
server's explicit re-attach), so a delete replay or stale repair can never
strip a live guest's NIC. Durably-bound ports are preserved by the network
cleanup fence until they unbind; stale-bound orphans of terminally-deleted
owners are swept via a deterministic unbind-then-release with fail-closed
retry. The new sweep-emitting lanes require process-wide REPORT_CAPTURE_LOCK
serialization during every emission window (the report buffer is
process-global), so all nine lock sites are a correctness prerequisite of
this commit's harness, identical to the final tree.
…1035)

Closes the sweep-unbind TOCTOU surfaced in adversarial review of the #1035
commit: a live server re-attaching an orphaned endpoint between the sweep's
durable-reference re-read and its port_binding unbind check could have its
realized binding torn down, and a create racing the sweep could durably
reference a port the sweep just released.

ComputeService gains orphan_repair_lock (tokio Mutex). The orphan repair
sweep holds it for its whole pass; the compat and native create adapters hold
the same lock across [existing-port validation/resolution -> durable intent
persist]. The lock is always taken FIRST, before store reads and before
projector/network calls, so the ordering cannot invert with the network
mutation locks. Under the lock each create-vs-sweep race collapses to exactly
one durable outcome: create wins (non-terminal server references the port,
which survives the sweep) or sweep wins (port released, create's existing-port
validation fails cleanly and leaves no durable reference). Regressions on both
lanes exercise the race and assert the settled invariant; no non-terminal
server ever references a deleted port.
Strengthen the PostgreSQL evidence and isolation before the S5 campaign:
the conformance PostgreSQL lane now fails closed (a configured O3K_DATABASE_URL
with an unreachable server fails the suite instead of silently skipping), and
the remaining p13_1 / p13_4 / p12_4 / audit / p13_f2_r1 ignored lanes pin the
PG adapter behavior. The pp4 endpoint-lifecycle PG lanes acquire a shared
advisory lock (o3k-shared-test-database) on the operator-owned test database
so they serialize against every other destructive PG test group even across
separate cargo-test processes. Effective-backend and property proofs are
recorded for the external/disposable provider model.
Add the fail-closed PP.5 campaign-evidence validation: a mutation/negative
coverage suite (scripts/validate_pp5_evidence.py + tests/pp5_evidence_validation.sh,
wired into CI) that proves the validator itself rejects malformed, unproven,
or DB-credential-leaking artifacts, plus a pp5-evidence-certify workflow, the
PP.5 execution plan and host-maintenance operations docs, and the small-edge
campaign host lifecycle scripts. Small-edge-cloud minimum target hypervisors
is corrected to 1. The CLI doctor test harness is hardened to fail closed
(panic on binary-execution failure instead of returning a synthetic exit
code that masks the cause).
…1036 #1037)

The external/disposable PostgreSQL provider mode, o3kd.env wiring with
effective-backend proof (pg_stat_activity pool count + proxy sever/restore),
and the six-identity five-concurrent BuildingBlock lifecycle are one campaign
artifact and are committed together: the journey evidence-generation python
builds a single evidence doc that fuses database_ownership and
scale_composition, so splitting them would leave two non-runnable halves.

The journey proves five identities concurrently Ready at peak and in the final
topology, with capacity baselines measured against the CURRENT topology before
each join (never self-referentially), drain-before-remove-before-replace
ordering, and a drained identity resolved canonically. The real-host and
fast-lane workflows authorize exactly block-a..block-f. The scale guard and
P15.7 protection suites (keycloak authority, dispatch, protected preflight,
postgres cleanup, libvirt pool ownership, real-host workflow guards) all pass.
Harden two PP.5 test flakes. The o3k-dhcp liveness assertion's injected
'sleep 0.2' delay could overrun the 500ms LIVENESS_GRACE window under heavy
parallel-test CPU load, flipping a legitimate setup-margin failure; the delay
is shrunk to 0.02s (well inside the grace) so the assertion tests the budget,
not scheduling noise. The compute create_convergence test re-arms its wait
budgets so one independent convergence wait does not consume the next one's
deadline, and separates the terminal-observation wait from the ACTIVE
projection wait it was sharing.
…hment (#1035)

Four fixes from adversarial review of the orphan-endpoint repair work:

1. HIGH - project_terminal_binding_outcome's delete branch (reachable via
   replayed agent terminal updates and in-request deletes) now gates every
   port on the live non-terminal reference set and holds the orphan-repair
   lock across [scan -> unbind -> release], so a late/replayed terminal update
   can no longer strip a port a NEW live server re-attached.
2. MEDIUM - every delete-replay seat and direct delete handler now holds the
   lock across [referenced-scan -> unbind -> release]: unbind_ports_from_intent
   and release_server_endpoints_from_intent lock internally, and the compat
   and native delete handlers acquire the orphan-repair lock for their direct
   network cleanup (acquired before the network-mutation lock).
3. MEDIUM-HIGH availability - the repair sweep now dispatches at most ONE
   fabric unbind per pass; remaining bound orphans are retried on the next
   5s pass, so a burst of stale-bound orphans cannot hold the repair lock (and
   block every port-attaching create) for the sum of dispatch deadlines. The
   honest worst case is one repair-path unbind deadline (~30s, matching the
   request-path fabric teardown) of create stall per pass while stale-bound
   orphans exist.
4. LOW - cleanup_server_owned_ports_for_project now holds the network-service
   lock across [binding-fence read -> delete] (neither sub-call re-locks), so
   the re-read and delete are atomic against a binding transition.

Re-entrancy audit: no caller of the fenced functions holds orphan_repair_lock
when the fenced scope runs. The create adapters hold the lock only across a
create (the delete branch is create-branch-excluded); the agent terminal
consumer, read surfaces, reconciler, and delete-server paths acquire it as one
layer. The API delete handlers narrow network_mutation_lock so orphan is always
acquired first.

Final lock graph: orphan_repair_lock (outermost) -> network_mutation_lock ->
NetworkService lock -> store. Regressions on both lanes: the replayed-terminal
delete projection preserves a live re-attached port, and the sweep converges
one unbind per pass.
@zoorpha
zoorpha requested a review from senolcolak as a code owner September 23, 2026 15:36
The maintainability SQL-boundary guardrail flagged the raw
sqlx::query('SELECT pg_advisory_lock(...)') in the pp4 endpoint-lifecycle
harness (o3kd composition), which is outside the approved persistence /
database-diagnostic / upgrade boundaries. Move that SQL into the persistence
crate (crates/o3k-store/src/conformance.rs, an approved boundary) as a small
public helper acquire_shared_postgres_test_database_lock, call it from the
pp4 advisory guard, and reuse it from the store's own
prepare_shared_postgres_test_database so there is exactly one SQL location.
Guardrails (maintainability + architecture) pass; pp4 SQLite (18) and PG
(23) lanes and the o3k-store PG conformance/terminalization lane stay green.
senolcolak
senolcolak previously approved these changes Sep 23, 2026
Drain-blocker derivation counted retained DELETED resource tombstones as
resident workload/local-storage/attachment blockers because
list_resources_by_kind deliberately returns tombstones and the adapter
applied no terminal filter. On PostgreSQL this became user-visible when
the store parity fix started returning tombstones; on SQLite it was
latent since the BuildingBlock drain-first contract landed. A block whose
workloads reached terminal DELETED could stay falsely drain-blocked and
unremovable.

The projection now skips observed_state == DELETED records, mirroring the
flavors terminal filter idiom, pinned by fail-before/fix-after adapter
regressions covering drain and remove semantics for tombstones and the
unchanged blocking behavior of live resources.
senolcolak
senolcolak previously approved these changes Sep 23, 2026
)

The agent event-stream seat terminalized a lifecycle delete with two
sequential durable writes (update_operation then the resource
projection), bypassing the terminalize_lifecycle primitive the
dispatch/poll finish path uses. A transient store failure between them
durably exposed a terminal operation whose DELETED projection never
landed, and nothing repairs it: the lifecycle sweep only re-drives
non-terminal operations and the observation-rejection path only logs.

A succeeded lifecycle:delete update now commits the terminal operation
and the DELETED resource projection in ONE transaction, provider
reference attach preserved ahead of it. Crash-arm and success/replay
regressions run against the TerminalizationFaultStore harness
(fail-before verified: the crash arm leaves the operation non-terminal
with the projection untouched).
… sweep (#1035)

Adversarial review found two MEDIUM gaps in the shipped repair sweep:

- The one-unbind-per-pass cap only counted successful dispatches. A
  failed unbind warned and continued to the next orphan, so a fabric
  outage let one pass dispatch N failing unbinds (each up to the ~30s
  fabric deadline) while holding the orphan-repair lock, violating the
  documented one-deadline availability bound. The cap now consumes the
  pass's single fabric budget on ATTEMPT and the pass ends; remaining
  bound orphans retry on the next periodic pass.
- The repair pass ran unconditionally on every controller. The
  orphan_repair_lock is process-local, so two controllers on a shared
  PostgreSQL database could sweep concurrently and degrade the
  create-validate/persist fence. The pass now acquires the same
  coordination work lease that fences the re-drive arms (fixed key
  'server-endpoint-orphan-repair', 60s bound); a Busy lease skips the
  pass and the lease is released after the pass.

Regressions (fail-before verified for the cap): at most one unbind
attempt per pass under persistent fabric failure with no release while
bound and cross-pass convergence; Busy lease skips the sweep, a free
lease repairs and is released.
A stale generation under a terminalizing write rolls the whole
transaction back and is an optimistic-concurrency conflict (a
concurrent observation legitimately bumped the row), but it fell into
the compute error catch-all as HTTP 500, misreporting the
exactly-one-winner fence as an outage. Map StoreError::StaleGeneration
to 409 Conflict, matching the revive path's mapping.
The external-PostgreSQL backend switch restarts o3kd and immediately
fires an authenticated 'o3k init' (only join has retries; init has
none). restart_o3kd_verified returned at process detection, but o3kd
binds AUTH_PORT only after pool init/seeding (~0.3s best case, ~1.5-2s
canonical on a fast idle host; the old process also takes ~5s to exit).
The init therefore deterministically raced the bind with
ECONNREFUSED on a fast host (latently flaky on slower CI). Reproduced
3x at dd5cf40 on the PP.5 S5 rerun; the journey never reached evidence
generation.

The restart now returns only after the control plane accepts HTTP on
AUTH_PORT (any response; /readyz stays deliberately gated behind the
rejoin). Guard suite re-verified.
The S5 journey at run 990923002 attempt 4 failed provisioning with
"SSH unavailable on real VM" against healthy guests. Live diagnostics
reproduced the full mechanism on the real host:

1. Stale-lease freeze: the journey derives MACs deterministically from
   the run id and dnsmasq retains leases for up to an hour, so a rerun
   observes the previous boot's lease. find_ip accepted the first lease
   it saw, froze that address for the whole SSH window, and the budget
   expired against an address the live guest did not hold.
2. Multi-lease first-match bias: each rerun mints a new DHCP client id,
   so dnsmasq accumulates one lease per boot per MAC; the stale entries
   sorted first, so naive first-match selection froze a stale address
   even after re-resolution.
3. pool-list parsing: libvirt 10.x pads 'pool-list --name' output with
   trailing spaces, so the helper's 'grep -Fxq' pool detection never
   matched on the real host; pool cleanup silently retained pools
   (the guard fixture emitted unpadded names, masking this).

Fix:
- New scripts/p15-7-vm-address.sh resolver: UUID-only lookups (never
  domain names), exact-MAC lease binding across domifaddr and
  net-dhcp-leases, gateway rejection, and freshest-lease selection via
  libvirt's dnsmasq status JSON (max expiry-time per MAC). When
  freshness data is unavailable it emits every MAC-bound candidate
  instead of guessing; SSH remains the only liveness proof.
- Journey wait_vm_ssh: single bounded 300x2s DHCP+SSH window
  (matches this host's documented multi-guest boot profile while
  keeping the six-VM journey inside the 120-min protected job budget),
  re-resolving the address every retry and liveness-probing each
  emitted candidate.
- Journey cleanup/assert_owned_domains_absent: prefer the recorded
  libvirt UUID as the locator once captured; the name is display-only.
- p15-7-libvirt-storage-pool.sh: strip libvirt 10.x column padding in
  pool_listing so pool detection/cleanup actually engages.
- Regressions: tests/p15_7_vm_address.sh (11 fake-virsh scenarios,
  incl. stale-first ordering, freshness selection, candidate-set
  fallback, UUID-only audit); pool cleanup fixture now emits the real
  padded name shape (fail-before demonstrated); guard assertions
  extended; CI step added.

restart_o3kd_verified TCP readiness retained on its own merits.

Co-authored-by: Kimi Code <noreply@moonshot.ai>
senolcolak
senolcolak previously approved these changes Sep 23, 2026
On run 990923002 the journey died silently immediately after the
Keycloak exchange: the unguarded single-shot
PROJECT_TOKEN="$(openstack token issue ...)" assignment hit a
transient control-plane failure (o3kd logged mass agent lease-renewal
database errors across all three connected agents at the same
instant, consistent with the run-owned PostgreSQL proxy stalling) and
`set -e` exited the whole journey with no diagnostics before the
existing die check could run. The identical authenticated read
succeeded minutes later against the same backend.

Fix: bounded 5-attempt retry with the loud die preserved on
persistent failure, matching the journey's existing retry conventions
(rejoin, agent readiness) and the `|| true` guard pattern already
used for the same CLI call at the cross-tenant probe. Token issue is
read-only; retrying a transient transport failure weakens no check.
Guard assertion pins the retry form.

Co-authored-by: Kimi Code <noreply@moonshot.ai>
senolcolak and others added 5 commits September 23, 2026 22:31
The bounded project-token retry added previously still exhausted all
five attempts within seconds on run 990923002 and the journey died with
only "authenticated project token unavailable", because the
openstack CLI's own stderr was discarded. Widen the bounded window to
15 attempts and capture the CLI error text to a run-owned file,
printing it (credential-shaped material redacted) on exhaustion so the
real cause -- status code, transport stall, or credential rejection --
is never guessed.

Co-authored-by: Kimi Code <noreply@moonshot.ai>
The journey read the placement host projection with the upper-case
spelling `OS-EXT-SRV-ATTR:HOST`, which the projection never returns:
the product serializes the field as `OS-EXT-SRV-ATTR:host` (matching
Nova) and the CLI resolves the column case-sensitively, so the host
came back empty for all 60 attempts and the drain/placement resolution
died with "workload A placement host has no unique ready canonical
block/provider mapping" even though the workload was ACTIVE and the
projection did carry the host (verified live: compute-agent). The
incorrect spelling had never been exercised because the journey had
not previously reached the workload stage on this host.

Guard assertion pins the lower-case contract and forbids the
upper-case spelling from returning.

Co-authored-by: Kimi Code <noreply@moonshot.ai>
The late PostgreSQL fault gate in external mode polled the
operator-owned server's own direct endpoint (O3K_DATABASE_URL), which
severing the run-owned proxy cannot affect: the loop always saw a
healthy endpoint and the journey died with "external PostgreSQL outage
was not observed" after 60 seconds. The gate could never pass in
external mode, which is why it had never been exercised (protected CI
runs use disposable mode).

Fix: inject the outage by severing the run-owned proxy (unchanged) and
observe the CONTROL PLANE's outage through the same signals the early
wiring proof uses -- /readyz gating and an authenticated durable-store
write -- then additionally assert the operator-owned server remains
reachable at its own endpoint (proving O3K never manages it), and
assert recovery after the proxy is restored. Guard assertion pins the
new signals and forbids the unreachable-endpoint polling loop.

Co-authored-by: Kimi Code <noreply@moonshot.ai>
On run 990923002 the journey completed the entire topology, drain/
remove/replace, restart, external PostgreSQL fault gate and SQLite
parity phases, then failed in the FINAL cleanup: the workload absence
proof hit a transient "compute service is unavailable" (HTTP 500) from
the control plane and the fail-closed classification retained owned
files, so assert_owned_domains_absent died on a retained seed ISO. The
same request returned a correct 404 minutes later against the same
backend, confirming a transient.

Fix: delete_owned_openstack retries its presence/delete/absence proof a
bounded number of times before declaring an object unprovable. The
fail-closed semantics are unchanged for genuinely unresolvable
outcomes; only transient control-plane failures are absorbed.

Co-authored-by: Kimi Code <noreply@moonshot.ai>
The sweep-cap regression added in 81e4cf1 left an unused
`o3k_store::DurableStore` import and used the now-lint-flagged
`io::Error::new(ErrorKind::Other, ..)` form, which the workspace clippy
gate (`-D warnings`, all targets) rejects. Use `io::Error::other` and
drop the unused import; behaviour of the regression is unchanged.

Co-authored-by: Kimi Code <noreply@moonhot.ai>
senolcolak
senolcolak previously approved these changes Sep 23, 2026

@senolcolak senolcolak left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

all green

senolcolak
senolcolak previously approved these changes Sep 27, 2026
senolcolak
senolcolak previously approved these changes Sep 27, 2026
senolcolak
senolcolak previously approved these changes Sep 27, 2026
senolcolak
senolcolak previously approved these changes Sep 28, 2026
senolcolak
senolcolak previously approved these changes Sep 28, 2026
senolcolak
senolcolak previously approved these changes Sep 28, 2026
senolcolak
senolcolak previously approved these changes Sep 28, 2026
senolcolak
senolcolak previously approved these changes Sep 28, 2026
@zoorpha

zoorpha commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator Author

Final protected #1035 attempt at 3d8009df (run 36493088830) failed before mutation: the targeted checkpoint was observed, then failed closed. No checkpoint, pre-mutation, release, unbind, or endpoint-release events followed; the target endpoint remained preserved, cleanup passed, and foreign state was unchanged.

The bounded artifacts cannot identify the publication step because failure_reason is not captured and the raw daemon log is not retained. Per the final-attempt rule, no further #1046 changes or protected retries will be made.

I am re-scoping #1046 as an uncertified PP.5 foundation and tracking publication-failure diagnosis separately in follow-up issue #1051.

This branch had an error being deployed

1 failed deployment
o3k-real-host-validation — 3d8009df Deployed Sep 29, 2026 by zoorpha via Protected TestLab real-host validation #471
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants