Platform
a5 (Ascend 950 hardware)
Runtime Variant
host_build_graph
Description
tests/st/run_retention/test_run_retention.py::test_result_reads_while_successor_runs_hbg intermittently fails on the current main baseline (818ee04d7308e129d7158ccbd8d078b53def4282) because its five attempts do not observe the successor still Pending after reading the predecessor's result. The predecessor result is valid on every failed attempt. The test uses a fixed 1,024-task successor workload to create an overlap window, so the failure currently does not distinguish a read that waited for the successor from a successor that naturally completed during the read.
This is a focused follow-up to the broad CI tracking issue. Related: #2425. PR #2388 encounters the same failure in CI, but its diff does not change either retention test file. The main-baseline failure below occurred before any PR-arm test was run in that device session; it does not establish that #2388 has zero effect on failure probability.
Steps to Reproduce
-
On an Ascend 950DT development device, build main commit 818ee04d7308e129d7158ccbd8d078b53def4282 with the A5 host_build_graph runtime and its pinned PTO-ISA. Run with the normal device lock and CANN environment.
-
Run the unmodified test repeatedly as independent pytest processes on the same device:
python -m pytest tests/st/run_retention/test_run_retention.py::test_result_reads_while_successor_runs_hbg --platform a5 --device 1 -v --require-pto-isa --pto-session-timeout 1200 --manual exclude
The reproduced runs used SIMPLER_SCHEDULER_TIMEOUT_MS=2000, SIMPLER_OP_EXECUTE_TIMEOUT_US=3000000, and SIMPLER_STREAM_SYNC_TIMEOUT_MS=4000.
-
Inspect the five-attempt assertion when the test fails.
On one fixed Ascend950DT device-1 session, the predeclared sequence was main 8 runs, #2388 8 runs, then main 8 runs. Main failed on runs 7 and 8 before the PR phase (6/8 passed); #2388 passed 5/8 and the final main phase passed 7/8. A separate fixed device-0 comparison passed main 5/8 and #2388 6/8. These small samples establish baseline reproducibility, not a reliable comparison of failure rates. Each arm used its own frozen host/AICPU/AICore/dispatcher build; the test source and orchestration .text were identical across arms. The device-1 task completed all 24 runs and released its lock (task_20260928_205820_175139424116).
Expected Behavior
The test should reliably establish that reading a completed predecessor's result does not wait for a still-running successor. If the intended overlap cannot be established, report that separately from a proven read-ordering violation; a genuinely blocked read must still fail.
Actual Behavior
At main commit 818ee04d, the test failed twice in eight consecutive independent processes before any #2388 run. A representative assertion was:
AssertionError: in 5 attempts the successor was never still Pending after the predecessor's read (attempts: ['completed-across-read', 'completed-across-read', 'no-overlap', 'completed-across-read', 'completed-across-read']). Every attempt's predecessor result was valid.
no-overlap means the successor completed before the read began. completed-across-read leaves two possibilities: the read waited, or the successor finished during the read. Current diagnostics do not distinguish them. PR #2388 CI runs 36401973884 and 36415323613 also failed this case. An earlier run 35832657012 passed it, so this is intermittent.
Git Commit ID
818ee04
CANN Version
9.2.0
Driver Version
Not captured for this reproduction.
Host Platform
Linux (aarch64)
Additional Context
The test and its 1,024-task window were added in PR #2330, before #2388 was opened. The baseline observation does not identify a regression-causing commit or prove a production-runtime defect. A useful fix should make the overlap condition observable or controlled, distinguish natural successor completion from a read that waits, and retain a negative check for actual read-ordering violations. Validate the unchanged main reproducer and relevant A5/A3 HBG retention coverage; do not rely only on increasing the fixed task count. The earlier 1,024-to-4,096-task proposal in PR #2469 was closed unmerged for discussion.
Platform
a5 (Ascend 950 hardware)
Runtime Variant
host_build_graph
Description
tests/st/run_retention/test_run_retention.py::test_result_reads_while_successor_runs_hbgintermittently fails on the current main baseline (818ee04d7308e129d7158ccbd8d078b53def4282) because its five attempts do not observe the successor stillPendingafter reading the predecessor's result. The predecessor result is valid on every failed attempt. The test uses a fixed 1,024-task successor workload to create an overlap window, so the failure currently does not distinguish a read that waited for the successor from a successor that naturally completed during the read.This is a focused follow-up to the broad CI tracking issue. Related: #2425. PR #2388 encounters the same failure in CI, but its diff does not change either retention test file. The main-baseline failure below occurred before any PR-arm test was run in that device session; it does not establish that #2388 has zero effect on failure probability.
Steps to Reproduce
On an Ascend 950DT development device, build main commit
818ee04d7308e129d7158ccbd8d078b53def4282with the A5host_build_graphruntime and its pinned PTO-ISA. Run with the normal device lock and CANN environment.Run the unmodified test repeatedly as independent pytest processes on the same device:
The reproduced runs used
SIMPLER_SCHEDULER_TIMEOUT_MS=2000,SIMPLER_OP_EXECUTE_TIMEOUT_US=3000000, andSIMPLER_STREAM_SYNC_TIMEOUT_MS=4000.Inspect the five-attempt assertion when the test fails.
On one fixed Ascend950DT device-1 session, the predeclared sequence was main 8 runs, #2388 8 runs, then main 8 runs. Main failed on runs 7 and 8 before the PR phase (6/8 passed); #2388 passed 5/8 and the final main phase passed 7/8. A separate fixed device-0 comparison passed main 5/8 and #2388 6/8. These small samples establish baseline reproducibility, not a reliable comparison of failure rates. Each arm used its own frozen host/AICPU/AICore/dispatcher build; the test source and orchestration
.textwere identical across arms. The device-1 task completed all 24 runs and released its lock (task_20260928_205820_175139424116).Expected Behavior
The test should reliably establish that reading a completed predecessor's result does not wait for a still-running successor. If the intended overlap cannot be established, report that separately from a proven read-ordering violation; a genuinely blocked read must still fail.
Actual Behavior
At main commit
818ee04d, the test failed twice in eight consecutive independent processes before any #2388 run. A representative assertion was:no-overlapmeans the successor completed before the read began.completed-across-readleaves two possibilities: the read waited, or the successor finished during the read. Current diagnostics do not distinguish them. PR #2388 CI runs 36401973884 and 36415323613 also failed this case. An earlier run 35832657012 passed it, so this is intermittent.Git Commit ID
818ee04
CANN Version
9.2.0
Driver Version
Not captured for this reproduction.
Host Platform
Linux (aarch64)
Additional Context
The test and its 1,024-task window were added in PR #2330, before #2388 was opened. The baseline observation does not identify a regression-causing commit or prove a production-runtime defect. A useful fix should make the overlap condition observable or controlled, distinguish natural successor completion from a read that waits, and retain a negative check for actual read-ordering violations. Validate the unchanged main reproducer and relevant A5/A3 HBG retention coverage; do not rely only on increasing the fixed task count. The earlier 1,024-to-4,096-task proposal in PR #2469 was closed unmerged for discussion.