Problem
The combined storage validation command can exit an S3 integration test with code 0 while Nextest still reports the
test process's captured stdout or stderr as open:
cargo nextest run --package peryx-storage --package peryx-ha-distributed --all-features --profile ci
The original run at f78ba1d7 completed 2,095 tests: 2,093 passed and two failed because of leaked handles. A diagnostic
run at the same commit reproduced LEAK-FAIL in a third S3 test while the original two passed. The affected test is
therefore not fixed to one multipart scenario.
Observed cases
Each stdout block ends with test ... ok, followed by test failed: exited with code 0, but leaked handles.
Diagnosis
Nextest 0.9.143 detects a leak
when the test process has exited but its captured stdout or stderr has not reached EOF within the configured second.
The count and executable hashes stayed stable across the diagnostic runs, which rules out stale shared-target binaries.
The S3 fixture cannot retain those Nextest pipes. The test launches it with Command::output; Tokio replaces
the fixture's stdout and stderr with private pipes
and waits for exit and EOF on both pipes.
Sampled passing runs found no reparented fixture descendant. The remaining diagnosis must identify every process or
duplicate file descriptor that retains the test process's actual Nextest pipe after exit.
A detached f78ba1d7 reuse run enumerated the same three executables and 2,095 tests without invoking Cargo's build
path. The HA and integration executables matched their original hashes. The shared target had overwritten the storage
library and S3 fixture, so those two files were rebuilt from the detached source but were not byte-identical to the
original run.
A temporary Nextest 0.9.143 diagnostic build emitted the runner's existing time_to_close_fds_nanos value without
changing leak detection. A low-overhead full run performed one relaxed atomic maximum update per test and
reported a suite maximum of 7,542 ns against the 1,000,000,000 ns limit. It passed all 2,095 tests and produced no
timeout event. The sampler captured 63 S3 test processes; none held a duplicate of its stdout or stderr pipe endpoint.
The passing run does not identify the historical owner. macOS reused pipe addresses for later processes, and Nextest
kept some read descriptors open after observing EOF, so neither an address match nor reader lifetime proves a retained
writer. The missing evidence is a failing timeout record containing the test identity, Nextest's reader descriptors,
and a simultaneous process-wide descriptor snapshot.
Required change
Capture the retaining PID and file descriptor during a failing run, then repair the owner that created it. Do not
change S3 fixture lifecycle code unless the capture proves that it owns the Nextest pipe.
Do not increase Nextest's leak timeout, exclude tests, disable leak detection, retry jobs, or add CI cleanup. Preserve
the multipart cleanup assertions and failure scenarios.
Acceptance checks
- A failing diagnostic names the process and file descriptor that retains the test's Nextest capture pipe.
- The owner closes and joins that resource before the test process exits.
- The combined command completes all three observed tests without
LEAK-FAIL.
- The completion-status failure still aborts the orphaned multipart upload and removes its journal.
- The concurrent-part case waits for started parts before sending one abort request.
- A regression fails when the repaired owner omits shutdown or join.
Architecture and test boundary
Keep S3 lifecycle behavior and its integration tests in peryx-storage. Repair the narrow owner that creates the
retained handle; do not add CI-specific cleanup logic or a repository script.
Problem
The combined storage validation command can exit an S3 integration test with code 0 while Nextest still reports the
test process's captured stdout or stderr as open:
The original run at
f78ba1d7completed 2,095 tests: 2,093 passed and two failed because of leaked handles. A diagnosticrun at the same commit reproduced
LEAK-FAILin a third S3 test while the original two passed. The affected test istherefore not fixed to one multipart scenario.
Observed cases
test_s3_cleans_up_when_peer_completion_status_cannot_be_checkedtest_s3_drains_started_parts_before_abortingtest_s3_reports_an_unreadable_multipart_journalEach stdout block ends with
test ... ok, followed bytest failed: exited with code 0, but leaked handles.Diagnosis
Nextest 0.9.143 detects a leak
when the test process has exited but its captured stdout or stderr has not reached EOF within the configured second.
The count and executable hashes stayed stable across the diagnostic runs, which rules out stale shared-target binaries.
The S3 fixture cannot retain those Nextest pipes. The test launches it with
Command::output; Tokio replacesthe fixture's stdout and stderr with private pipes
and waits for exit and EOF on both pipes.
Sampled passing runs found no reparented fixture descendant. The remaining diagnosis must identify every process or
duplicate file descriptor that retains the test process's actual Nextest pipe after exit.
A detached
f78ba1d7reuse run enumerated the same three executables and 2,095 tests without invoking Cargo's buildpath. The HA and integration executables matched their original hashes. The shared target had overwritten the storage
library and S3 fixture, so those two files were rebuilt from the detached source but were not byte-identical to the
original run.
A temporary Nextest 0.9.143 diagnostic build emitted the runner's existing
time_to_close_fds_nanosvalue withoutchanging leak detection. A low-overhead full run performed one relaxed atomic maximum update per test and
reported a suite maximum of 7,542 ns against the 1,000,000,000 ns limit. It passed all 2,095 tests and produced no
timeout event. The sampler captured 63 S3 test processes; none held a duplicate of its stdout or stderr pipe endpoint.
The passing run does not identify the historical owner. macOS reused pipe addresses for later processes, and Nextest
kept some read descriptors open after observing EOF, so neither an address match nor reader lifetime proves a retained
writer. The missing evidence is a failing timeout record containing the test identity, Nextest's reader descriptors,
and a simultaneous process-wide descriptor snapshot.
Required change
Capture the retaining PID and file descriptor during a failing run, then repair the owner that created it. Do not
change S3 fixture lifecycle code unless the capture proves that it owns the Nextest pipe.
Do not increase Nextest's leak timeout, exclude tests, disable leak detection, retry jobs, or add CI cleanup. Preserve
the multipart cleanup assertions and failure scenarios.
Acceptance checks
LEAK-FAIL.Architecture and test boundary
Keep S3 lifecycle behavior and its integration tests in
peryx-storage. Repair the narrow owner that creates theretained handle; do not add CI-specific cleanup logic or a repository script.