Conversation
Lifecycle success previously committed the operation terminal row and the resource terminal projection as two sequential durable writes; a crash between them left exactly one half committed, which no reconciliation pass repairs. terminalize_lifecycle now commits the operation REVERSIBLE->FINAL transition, the canonical metadata row, and the fenced (generation=expected) terminal resource projection in ONE transaction on both adapters (SQLite BEGIN IMMEDIATE; PG pool.begin + FOR UPDATE), with an equivalent replay as a no-op and a conflicting terminal -> Corrupt. This commit also makes the tombstone-visibility contract explicit (:89) and fixes the PostgreSQL adapter that violated it: list_resources_by_kind returns every resource of the kind including DELETED tombstones on both adapters, so repair scans (which exist precisely to reconcile tombstones) see durable terminal truth. Endpoint/network release stays outside the transaction, repaired by the orphan-endpoint sweep (#1035).
Ownership and repair authority for o3k-server endpoints. The request-path releases (native + compat delete handlers) and the new repair-only sweep now re-read the live-attached endpoint set immediately before releasing and refuse to hand back any port the compute layer still references (a NEW live server's explicit re-attach), so a delete replay or stale repair can never strip a live guest's NIC. Durably-bound ports are preserved by the network cleanup fence until they unbind; stale-bound orphans of terminally-deleted owners are swept via a deterministic unbind-then-release with fail-closed retry. The new sweep-emitting lanes require process-wide REPORT_CAPTURE_LOCK serialization during every emission window (the report buffer is process-global), so all nine lock sites are a correctness prerequisite of this commit's harness, identical to the final tree.
…1035) Closes the sweep-unbind TOCTOU surfaced in adversarial review of the #1035 commit: a live server re-attaching an orphaned endpoint between the sweep's durable-reference re-read and its port_binding unbind check could have its realized binding torn down, and a create racing the sweep could durably reference a port the sweep just released. ComputeService gains orphan_repair_lock (tokio Mutex). The orphan repair sweep holds it for its whole pass; the compat and native create adapters hold the same lock across [existing-port validation/resolution -> durable intent persist]. The lock is always taken FIRST, before store reads and before projector/network calls, so the ordering cannot invert with the network mutation locks. Under the lock each create-vs-sweep race collapses to exactly one durable outcome: create wins (non-terminal server references the port, which survives the sweep) or sweep wins (port released, create's existing-port validation fails cleanly and leaves no durable reference). Regressions on both lanes exercise the race and assert the settled invariant; no non-terminal server ever references a deleted port.
Strengthen the PostgreSQL evidence and isolation before the S5 campaign: the conformance PostgreSQL lane now fails closed (a configured O3K_DATABASE_URL with an unreachable server fails the suite instead of silently skipping), and the remaining p13_1 / p13_4 / p12_4 / audit / p13_f2_r1 ignored lanes pin the PG adapter behavior. The pp4 endpoint-lifecycle PG lanes acquire a shared advisory lock (o3k-shared-test-database) on the operator-owned test database so they serialize against every other destructive PG test group even across separate cargo-test processes. Effective-backend and property proofs are recorded for the external/disposable provider model.
Add the fail-closed PP.5 campaign-evidence validation: a mutation/negative coverage suite (scripts/validate_pp5_evidence.py + tests/pp5_evidence_validation.sh, wired into CI) that proves the validator itself rejects malformed, unproven, or DB-credential-leaking artifacts, plus a pp5-evidence-certify workflow, the PP.5 execution plan and host-maintenance operations docs, and the small-edge campaign host lifecycle scripts. Small-edge-cloud minimum target hypervisors is corrected to 1. The CLI doctor test harness is hardened to fail closed (panic on binary-execution failure instead of returning a synthetic exit code that masks the cause).
…1036 #1037) The external/disposable PostgreSQL provider mode, o3kd.env wiring with effective-backend proof (pg_stat_activity pool count + proxy sever/restore), and the six-identity five-concurrent BuildingBlock lifecycle are one campaign artifact and are committed together: the journey evidence-generation python builds a single evidence doc that fuses database_ownership and scale_composition, so splitting them would leave two non-runnable halves. The journey proves five identities concurrently Ready at peak and in the final topology, with capacity baselines measured against the CURRENT topology before each join (never self-referentially), drain-before-remove-before-replace ordering, and a drained identity resolved canonically. The real-host and fast-lane workflows authorize exactly block-a..block-f. The scale guard and P15.7 protection suites (keycloak authority, dispatch, protected preflight, postgres cleanup, libvirt pool ownership, real-host workflow guards) all pass.
Harden two PP.5 test flakes. The o3k-dhcp liveness assertion's injected 'sleep 0.2' delay could overrun the 500ms LIVENESS_GRACE window under heavy parallel-test CPU load, flipping a legitimate setup-margin failure; the delay is shrunk to 0.02s (well inside the grace) so the assertion tests the budget, not scheduling noise. The compute create_convergence test re-arms its wait budgets so one independent convergence wait does not consume the next one's deadline, and separates the terminal-observation wait from the ACTIVE projection wait it was sharing.
…hment (#1035) Four fixes from adversarial review of the orphan-endpoint repair work: 1. HIGH - project_terminal_binding_outcome's delete branch (reachable via replayed agent terminal updates and in-request deletes) now gates every port on the live non-terminal reference set and holds the orphan-repair lock across [scan -> unbind -> release], so a late/replayed terminal update can no longer strip a port a NEW live server re-attached. 2. MEDIUM - every delete-replay seat and direct delete handler now holds the lock across [referenced-scan -> unbind -> release]: unbind_ports_from_intent and release_server_endpoints_from_intent lock internally, and the compat and native delete handlers acquire the orphan-repair lock for their direct network cleanup (acquired before the network-mutation lock). 3. MEDIUM-HIGH availability - the repair sweep now dispatches at most ONE fabric unbind per pass; remaining bound orphans are retried on the next 5s pass, so a burst of stale-bound orphans cannot hold the repair lock (and block every port-attaching create) for the sum of dispatch deadlines. The honest worst case is one repair-path unbind deadline (~30s, matching the request-path fabric teardown) of create stall per pass while stale-bound orphans exist. 4. LOW - cleanup_server_owned_ports_for_project now holds the network-service lock across [binding-fence read -> delete] (neither sub-call re-locks), so the re-read and delete are atomic against a binding transition. Re-entrancy audit: no caller of the fenced functions holds orphan_repair_lock when the fenced scope runs. The create adapters hold the lock only across a create (the delete branch is create-branch-excluded); the agent terminal consumer, read surfaces, reconciler, and delete-server paths acquire it as one layer. The API delete handlers narrow network_mutation_lock so orphan is always acquired first. Final lock graph: orphan_repair_lock (outermost) -> network_mutation_lock -> NetworkService lock -> store. Regressions on both lanes: the replayed-terminal delete projection preserves a live re-attached port, and the sweep converges one unbind per pass.
The maintainability SQL-boundary guardrail flagged the raw
sqlx::query('SELECT pg_advisory_lock(...)') in the pp4 endpoint-lifecycle
harness (o3kd composition), which is outside the approved persistence /
database-diagnostic / upgrade boundaries. Move that SQL into the persistence
crate (crates/o3k-store/src/conformance.rs, an approved boundary) as a small
public helper acquire_shared_postgres_test_database_lock, call it from the
pp4 advisory guard, and reuse it from the store's own
prepare_shared_postgres_test_database so there is exactly one SQL location.
Guardrails (maintainability + architecture) pass; pp4 SQLite (18) and PG
(23) lanes and the o3k-store PG conformance/terminalization lane stay green.
senolcolak
previously approved these changes
Sep 23, 2026
Drain-blocker derivation counted retained DELETED resource tombstones as resident workload/local-storage/attachment blockers because list_resources_by_kind deliberately returns tombstones and the adapter applied no terminal filter. On PostgreSQL this became user-visible when the store parity fix started returning tombstones; on SQLite it was latent since the BuildingBlock drain-first contract landed. A block whose workloads reached terminal DELETED could stay falsely drain-blocked and unremovable. The projection now skips observed_state == DELETED records, mirroring the flavors terminal filter idiom, pinned by fail-before/fix-after adapter regressions covering drain and remove semantics for tombstones and the unchanged blocking behavior of live resources.
senolcolak
previously approved these changes
Sep 23, 2026
) The agent event-stream seat terminalized a lifecycle delete with two sequential durable writes (update_operation then the resource projection), bypassing the terminalize_lifecycle primitive the dispatch/poll finish path uses. A transient store failure between them durably exposed a terminal operation whose DELETED projection never landed, and nothing repairs it: the lifecycle sweep only re-drives non-terminal operations and the observation-rejection path only logs. A succeeded lifecycle:delete update now commits the terminal operation and the DELETED resource projection in ONE transaction, provider reference attach preserved ahead of it. Crash-arm and success/replay regressions run against the TerminalizationFaultStore harness (fail-before verified: the crash arm leaves the operation non-terminal with the projection untouched).
… sweep (#1035) Adversarial review found two MEDIUM gaps in the shipped repair sweep: - The one-unbind-per-pass cap only counted successful dispatches. A failed unbind warned and continued to the next orphan, so a fabric outage let one pass dispatch N failing unbinds (each up to the ~30s fabric deadline) while holding the orphan-repair lock, violating the documented one-deadline availability bound. The cap now consumes the pass's single fabric budget on ATTEMPT and the pass ends; remaining bound orphans retry on the next periodic pass. - The repair pass ran unconditionally on every controller. The orphan_repair_lock is process-local, so two controllers on a shared PostgreSQL database could sweep concurrently and degrade the create-validate/persist fence. The pass now acquires the same coordination work lease that fences the re-drive arms (fixed key 'server-endpoint-orphan-repair', 60s bound); a Busy lease skips the pass and the lease is released after the pass. Regressions (fail-before verified for the cap): at most one unbind attempt per pass under persistent fabric failure with no release while bound and cross-pass convergence; Busy lease skips the sweep, a free lease repairs and is released.
A stale generation under a terminalizing write rolls the whole transaction back and is an optimistic-concurrency conflict (a concurrent observation legitimately bumped the row), but it fell into the compute error catch-all as HTTP 500, misreporting the exactly-one-winner fence as an outage. Map StoreError::StaleGeneration to 409 Conflict, matching the revive path's mapping.
The external-PostgreSQL backend switch restarts o3kd and immediately fires an authenticated 'o3k init' (only join has retries; init has none). restart_o3kd_verified returned at process detection, but o3kd binds AUTH_PORT only after pool init/seeding (~0.3s best case, ~1.5-2s canonical on a fast idle host; the old process also takes ~5s to exit). The init therefore deterministically raced the bind with ECONNREFUSED on a fast host (latently flaky on slower CI). Reproduced 3x at dd5cf40 on the PP.5 S5 rerun; the journey never reached evidence generation. The restart now returns only after the control plane accepts HTTP on AUTH_PORT (any response; /readyz stays deliberately gated behind the rejoin). Guard suite re-verified.
The S5 journey at run 990923002 attempt 4 failed provisioning with "SSH unavailable on real VM" against healthy guests. Live diagnostics reproduced the full mechanism on the real host: 1. Stale-lease freeze: the journey derives MACs deterministically from the run id and dnsmasq retains leases for up to an hour, so a rerun observes the previous boot's lease. find_ip accepted the first lease it saw, froze that address for the whole SSH window, and the budget expired against an address the live guest did not hold. 2. Multi-lease first-match bias: each rerun mints a new DHCP client id, so dnsmasq accumulates one lease per boot per MAC; the stale entries sorted first, so naive first-match selection froze a stale address even after re-resolution. 3. pool-list parsing: libvirt 10.x pads 'pool-list --name' output with trailing spaces, so the helper's 'grep -Fxq' pool detection never matched on the real host; pool cleanup silently retained pools (the guard fixture emitted unpadded names, masking this). Fix: - New scripts/p15-7-vm-address.sh resolver: UUID-only lookups (never domain names), exact-MAC lease binding across domifaddr and net-dhcp-leases, gateway rejection, and freshest-lease selection via libvirt's dnsmasq status JSON (max expiry-time per MAC). When freshness data is unavailable it emits every MAC-bound candidate instead of guessing; SSH remains the only liveness proof. - Journey wait_vm_ssh: single bounded 300x2s DHCP+SSH window (matches this host's documented multi-guest boot profile while keeping the six-VM journey inside the 120-min protected job budget), re-resolving the address every retry and liveness-probing each emitted candidate. - Journey cleanup/assert_owned_domains_absent: prefer the recorded libvirt UUID as the locator once captured; the name is display-only. - p15-7-libvirt-storage-pool.sh: strip libvirt 10.x column padding in pool_listing so pool detection/cleanup actually engages. - Regressions: tests/p15_7_vm_address.sh (11 fake-virsh scenarios, incl. stale-first ordering, freshness selection, candidate-set fallback, UUID-only audit); pool cleanup fixture now emits the real padded name shape (fail-before demonstrated); guard assertions extended; CI step added. restart_o3kd_verified TCP readiness retained on its own merits. Co-authored-by: Kimi Code <noreply@moonshot.ai>
senolcolak
previously approved these changes
Sep 23, 2026
On run 990923002 the journey died silently immediately after the Keycloak exchange: the unguarded single-shot PROJECT_TOKEN="$(openstack token issue ...)" assignment hit a transient control-plane failure (o3kd logged mass agent lease-renewal database errors across all three connected agents at the same instant, consistent with the run-owned PostgreSQL proxy stalling) and `set -e` exited the whole journey with no diagnostics before the existing die check could run. The identical authenticated read succeeded minutes later against the same backend. Fix: bounded 5-attempt retry with the loud die preserved on persistent failure, matching the journey's existing retry conventions (rejoin, agent readiness) and the `|| true` guard pattern already used for the same CLI call at the cross-tenant probe. Token issue is read-only; retrying a transient transport failure weakens no check. Guard assertion pins the retry form. Co-authored-by: Kimi Code <noreply@moonshot.ai>
The bounded project-token retry added previously still exhausted all five attempts within seconds on run 990923002 and the journey died with only "authenticated project token unavailable", because the openstack CLI's own stderr was discarded. Widen the bounded window to 15 attempts and capture the CLI error text to a run-owned file, printing it (credential-shaped material redacted) on exhaustion so the real cause -- status code, transport stall, or credential rejection -- is never guessed. Co-authored-by: Kimi Code <noreply@moonshot.ai>
The journey read the placement host projection with the upper-case spelling `OS-EXT-SRV-ATTR:HOST`, which the projection never returns: the product serializes the field as `OS-EXT-SRV-ATTR:host` (matching Nova) and the CLI resolves the column case-sensitively, so the host came back empty for all 60 attempts and the drain/placement resolution died with "workload A placement host has no unique ready canonical block/provider mapping" even though the workload was ACTIVE and the projection did carry the host (verified live: compute-agent). The incorrect spelling had never been exercised because the journey had not previously reached the workload stage on this host. Guard assertion pins the lower-case contract and forbids the upper-case spelling from returning. Co-authored-by: Kimi Code <noreply@moonshot.ai>
The late PostgreSQL fault gate in external mode polled the operator-owned server's own direct endpoint (O3K_DATABASE_URL), which severing the run-owned proxy cannot affect: the loop always saw a healthy endpoint and the journey died with "external PostgreSQL outage was not observed" after 60 seconds. The gate could never pass in external mode, which is why it had never been exercised (protected CI runs use disposable mode). Fix: inject the outage by severing the run-owned proxy (unchanged) and observe the CONTROL PLANE's outage through the same signals the early wiring proof uses -- /readyz gating and an authenticated durable-store write -- then additionally assert the operator-owned server remains reachable at its own endpoint (proving O3K never manages it), and assert recovery after the proxy is restored. Guard assertion pins the new signals and forbids the unreachable-endpoint polling loop. Co-authored-by: Kimi Code <noreply@moonshot.ai>
On run 990923002 the journey completed the entire topology, drain/ remove/replace, restart, external PostgreSQL fault gate and SQLite parity phases, then failed in the FINAL cleanup: the workload absence proof hit a transient "compute service is unavailable" (HTTP 500) from the control plane and the fail-closed classification retained owned files, so assert_owned_domains_absent died on a retained seed ISO. The same request returned a correct 404 minutes later against the same backend, confirming a transient. Fix: delete_owned_openstack retries its presence/delete/absence proof a bounded number of times before declaring an object unprovable. The fail-closed semantics are unchanged for genuinely unresolvable outcomes; only transient control-plane failures are absorbed. Co-authored-by: Kimi Code <noreply@moonshot.ai>
The sweep-cap regression added in 81e4cf1 left an unused `o3k_store::DurableStore` import and used the now-lint-flagged `io::Error::new(ErrorKind::Other, ..)` form, which the workspace clippy gate (`-D warnings`, all targets) rejects. Use `io::Error::other` and drop the unused import; behaviour of the regression is unchanged. Co-authored-by: Kimi Code <noreply@moonhot.ai>
senolcolak
previously approved these changes
Sep 27, 2026
senolcolak
previously approved these changes
Sep 27, 2026
senolcolak
previously approved these changes
Sep 27, 2026
senolcolak
previously approved these changes
Sep 28, 2026
senolcolak
previously approved these changes
Sep 28, 2026
senolcolak
previously approved these changes
Sep 28, 2026
senolcolak
previously approved these changes
Sep 28, 2026
senolcolak
previously approved these changes
Sep 28, 2026
Collaborator
Author
|
Final protected #1035 attempt at The bounded artifacts cannot identify the publication step because I am re-scoping #1046 as an uncertified PP.5 foundation and tracking publication-failure diagnosis separately in follow-up issue #1051. |
This was referenced Sep 29, 2026
senolcolak
approved these changes
Sep 29, 2026
This branch had an error being deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
Build the reviewed PP.5 commit stack on the
pp5-s5-foundationbranch: 8 commits covering the five-concurrent BuildingBlock lifecycle campaign, the PostgreSQL provider/isolation hardening, fail-closed campaign evidence gates, and the deterministic-timing sweep.Scope
terminalize_lifecyclesingle-transaction primitive on SQLite + PostgreSQL; generation fence in-tx; replay no-op) and the tombstone-visibility contract (list_resources_by_kindreturns DELETED tombstones on both adapters).Files: 45 changed, ~6800 insertions.
Product / deployment profile
small-edge-cloud(minimum target hypervisors corrected to 1).Non-goals (explicit)
OpenStack compatibility impact
/v2.1//v2.0contract edges in the pp4/pp5 harness; no new advertised OpenStack service surface. O3K internal architecture (shared Cloud Kernel) is unchanged.Database / footprint
operations/resourcesrows. PostgreSQL PG + SQLite lanes both green.Authority model / workflow / compensation
orphan_repair_lock(leaf-first ordering: orphan_repair_lock → network_mutation_lock → NetworkService lock → store). Replayed/late terminal updates and delete replays are fenced against live re-attachment; deterministico3k:orphan-unbindop id with fail-closed retry.Evidence / tests
Run locally on
e03223df(RTK-disabled, real PostgreSQL viaO3K_DATABASE_URL):cargo fmt --all -- --check— passcargo clippy --workspace --all-targets --all-features -- -D warnings— passcargo test --workspace --all-features --no-fail-fast(with PG) — 1748 passed / 0 failedpp5_evidence_validation.sh)Adversarial review
The endpoint-repair commits received adversarial review against TOCTOU / stale-binding / cross-project-identity / re-attachment / sweep-capacity attack vectors. Two findings were fixed and re-reviewed clean over commits
fce0adc5→9b3ff63d(create-vs-sweep serialization) and →e03223df(fence every release seat; sweep cap; atomic binding fence). (Note from executor: one independent reviewer could not be spawned in my subagent environment — the review was performed with the same structured methodology and is recorded in the commit bodies.)Request
Review requested; attention commits:
fce0adc5,9b3ff63d,e03223df.