Summary
simpler PR hw-native-sys/simpler#2488 only updated the pinned PTO-ISA revision from c0d7148e95ef73bd12a73165fdce4b723a3b7e72 to 15a9e0a0845955f5d7a409a7f4d1609a263b5d25. Both subsequent st-network1-onboard-a2a3 CI runs failed. The complete device logs (plog) from the second run confirm an AIV vector instruction accessing Unified Buffer (UB) memory out of bounds, followed by scheduler timeouts and worker cleanup errors.
This is a suspected regression supported by an old-pin control run on the same runner and devices. A causal link to a specific PTO-ISA commit has not been established. This issue tracks endpoint revalidation and bisection.
Observed behavior
| Test case |
New pin: first run |
New pin: second run |
Old-pin control: immediately subsequent run |
compute_then_tload_mixed_l3 |
FAIL |
PASS |
PASS |
global_tload_mixed_l3 |
PASS |
FAIL |
PASS |
global_tload_mpirun_l3 |
PASS |
FAIL |
PASS |
vector_add_mixed_l3 |
PASS |
FAIL |
PASS |
global_tload_mixed_l3_network1 |
PASS |
PASS |
PASS |
The second run reported the following error sequence:
fftsplus aivector error
error code = 0x4000000000000000
The address for the VEC instruction to read/write UB is out of bounds.
orch_error_code=0 sched_error_code=100 runtime_status=-100
scheduler timeout sub_class=S1:running-stalled
aclrtSynchronizeStreamWithTimeout (AICPU) failed: 507018
global_tload_mixed_l3 and global_tload_mpirun_l3 failed on device 13 of the local node, with completed=0/1, stuck_task_id=0x0, and stuck_core=26. vector_add_mixed_l3 failed on local devices 12 and 13, with completed=2/5, stuck_task_id=0x100000001, and kernel ID 1 (kernel_add_scalar). The failure is therefore not limited to cross-node peer-memory TLOAD operations. The corresponding tasks on the peer node completed successfully.
Expected behavior
All five network1 test cases should continue to pass after the pin update, without VEC UB out-of-bounds errors, scheduler timeouts, or subsequent cleanup failures.
Reproduction procedure and prerequisites
Use simpler source revision 370c6633e05693020ef73adf97cfbeede93698b1 and the existing two-node network1 CI environment. Rebuild and reinstall on both nodes with the PTO-ISA revision under test. Preserve the existing device isolation and timeout settings, and run the network1 workflow at that simpler revision.
The workflow uses the following test invocation:
python -m pytest examples tests/st --level 4 --platform a2a3 \
--device <reserved-local-devices> --max-parallel 1 -v \
--require-pto-isa --pto-session-timeout 1200
This invocation requires the peer daemon, MPI launcher, and devices on both nodes configured by the workflow. It is not a standalone reproducer without those prerequisites.
Before starting bisection, repeat an old-pin / new-pin / old-pin comparison with the simpler source, compiler, and devices held constant. Start bisection only after the endpoints can be distinguished reliably. The two failing runs affected different sets of test cases, so a single passing run is insufficient to classify a candidate as good. Unrelated initialization failures must not be classified as bad for this regression.
Environment and versions
- Platform: Linux ARM64; self-hosted runner
a2a3pod-1; A2/A3 onboard two-node CI lane.
- Devices: local devices 12 and 13; peer devices 14 and 15. The exact SoC model and the CANN/ccec versions used in CI still need to be collected by a diagnostic job. Versions from the investigation host must not be used as substitutes.
- simpler runtime:
tensormap_and_ringbuffer.
- Previous PTO-ISA revision:
c0d7148e95ef73bd12a73165fdce4b723a3b7e72.
- Updated PTO-ISA revision:
15a9e0a0845955f5d7a409a7f4d1609a263b5d25 (the comparison range contains 49 commits).
- simpler merge commit for the first CI run:
932867b846b05470dfda8227f5e9c2cac6b44770.
- simpler merge commit for the second CI run:
370c6633e05693020ef73adf97cfbeede93698b1.
- simpler merge commit for the old-pin control:
6b5850bce9628059c52c05c60ed7d50640d25e48.
- Timeout settings:
SIMPLER_SCHEDULER_TIMEOUT_MS=2000, SIMPLER_OP_EXECUTE_TIMEOUT_US=3000000, and SIMPLER_STREAM_SYNC_TIMEOUT_MS=4000.
Verified evidence and limitations
- The second failing run and the immediately subsequent old-pin control used the same runner and devices. A full Git tree comparison confirmed that their A2/A3 runtime code, network1 test cases, and workflow files were identical. Apart from the pin update, the differences were confined to A5, documentation, and one Python unit test file.
- The local AICPU images in both jobs had the same fingerprint,
7c8cba0c0254fc4e, and the same size, 4173384 bytes. The actual user kernel ELF files and loaded code have not been uploaded, so this does not establish that the user kernels were identical.
- On the investigation host with CANN 9.0, the loaded
.text sections of the five relevant kernels and the TMR runtime files were identical between the old and new pins. This does not establish that the CI build artifacts were identical. Hardware testing on the investigation host was blocked by a separate bootstrap initialization error, so it did not produce a valid runtime good/bad comparison.
- A nearby old-pin job also reported
simpler_init ... 507033. Its failure stage and error signature differ from those of this issue, so it must not be treated as a reproduction of the UB out-of-bounds error.
- The first run's artifacts did not include local device logs for the failing test case. The second run's artifact,
network1-ci-36663462322-1, includes the relevant plog files and scheduler snapshots.
Related issue and next steps
Related: #170 also concerns A2/A3 UB behavior, but reports NaN output near the UB capacity limit. There is currently no evidence that the two issues share a root cause.
Bisection results will be added to this issue. No specific upstream commit will be identified as the root cause before endpoint revalidation is complete.
Summary
simpler PR hw-native-sys/simpler#2488 only updated the pinned PTO-ISA revision from
c0d7148e95ef73bd12a73165fdce4b723a3b7e72to15a9e0a0845955f5d7a409a7f4d1609a263b5d25. Both subsequentst-network1-onboard-a2a3CI runs failed. The complete device logs (plog) from the second run confirm an AIV vector instruction accessing Unified Buffer (UB) memory out of bounds, followed by scheduler timeouts and worker cleanup errors.This is a suspected regression supported by an old-pin control run on the same runner and devices. A causal link to a specific PTO-ISA commit has not been established. This issue tracks endpoint revalidation and bisection.
Observed behavior
compute_then_tload_mixed_l3global_tload_mixed_l3global_tload_mpirun_l3vector_add_mixed_l3global_tload_mixed_l3_network1The second run reported the following error sequence:
global_tload_mixed_l3andglobal_tload_mpirun_l3failed on device 13 of the local node, withcompleted=0/1,stuck_task_id=0x0, andstuck_core=26.vector_add_mixed_l3failed on local devices 12 and 13, withcompleted=2/5,stuck_task_id=0x100000001, and kernel ID 1 (kernel_add_scalar). The failure is therefore not limited to cross-node peer-memory TLOAD operations. The corresponding tasks on the peer node completed successfully.Expected behavior
All five network1 test cases should continue to pass after the pin update, without VEC UB out-of-bounds errors, scheduler timeouts, or subsequent cleanup failures.
Reproduction procedure and prerequisites
Use simpler source revision
370c6633e05693020ef73adf97cfbeede93698b1and the existing two-node network1 CI environment. Rebuild and reinstall on both nodes with the PTO-ISA revision under test. Preserve the existing device isolation and timeout settings, and run the network1 workflow at that simpler revision.The workflow uses the following test invocation:
This invocation requires the peer daemon, MPI launcher, and devices on both nodes configured by the workflow. It is not a standalone reproducer without those prerequisites.
Before starting bisection, repeat an old-pin / new-pin / old-pin comparison with the simpler source, compiler, and devices held constant. Start bisection only after the endpoints can be distinguished reliably. The two failing runs affected different sets of test cases, so a single passing run is insufficient to classify a candidate as good. Unrelated initialization failures must not be classified as bad for this regression.
Environment and versions
a2a3pod-1; A2/A3 onboard two-node CI lane.tensormap_and_ringbuffer.c0d7148e95ef73bd12a73165fdce4b723a3b7e72.15a9e0a0845955f5d7a409a7f4d1609a263b5d25(the comparison range contains 49 commits).932867b846b05470dfda8227f5e9c2cac6b44770.370c6633e05693020ef73adf97cfbeede93698b1.6b5850bce9628059c52c05c60ed7d50640d25e48.SIMPLER_SCHEDULER_TIMEOUT_MS=2000,SIMPLER_OP_EXECUTE_TIMEOUT_US=3000000, andSIMPLER_STREAM_SYNC_TIMEOUT_MS=4000.Verified evidence and limitations
7c8cba0c0254fc4e, and the same size, 4173384 bytes. The actual user kernel ELF files and loaded code have not been uploaded, so this does not establish that the user kernels were identical..textsections of the five relevant kernels and the TMR runtime files were identical between the old and new pins. This does not establish that the CI build artifacts were identical. Hardware testing on the investigation host was blocked by a separate bootstrap initialization error, so it did not produce a valid runtime good/bad comparison.simpler_init ... 507033. Its failure stage and error signature differ from those of this issue, so it must not be treated as a reproduction of the UB out-of-bounds error.network1-ci-36663462322-1, includes the relevantplogfiles and scheduler snapshots.Related issue and next steps
Related: #170 also concerns A2/A3 UB behavior, but reports NaN output near the UB capacity limit. There is currently no evidence that the two issues share a root cause.
Bisection results will be added to this issue. No specific upstream commit will be identified as the root cause before endpoint revalidation is complete.