Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
227 changes: 227 additions & 0 deletions docs/investigations/2026-09-a5-aicore-retirement-port.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,227 @@
# 2026-09 — Porting the a2a3 AICore retirement protocol to a5: what was dropped on the way

**Verdict: preserve a2a3's per-core claim, emergency batching and post-close
return gate.** PR #2388 (issue #2387) ports these to a5 TMR and covers the
corresponding HBG teardown paths. Separate fast-path
register control has no defined a5 software counterpart. Historical measurements
below describe earlier candidates, not the current per-core gated implementation.

Supersedes the verdict of the entry proposed in #2300 ("dropped, no net gain"),
which was measured over a shape this one does not use. See *Where the earlier
verdict went wrong* below.

## Setting

a5 retired AICores one at a time: write EXIT, block for that core's ACK
against a per-core deadline, write IDLE. a2a3 retires the same hardware as a
claimed group with a read-back close. The question was which of a2a3's
mechanisms apply to a5, and what they cost.

Measurements: a5 development card, nine graphs by three blocks, paired within
each block, medians of 20 timed calls per process, five rounds of 171
processes. The segment is the retirement tail — `graph_build` end minus
`max(orch end, sched end)` — derived from the `[STRACE]` markers already
emitted. Parenthesised counts are blocks where the sign held.

## Protocol

The mechanisms are the per-core claim, `wmb` after the broadcast, a non-blocking
round-robin sweep against one shared deadline, deferred window close, the
close read-back, the `rmb` that drains it, per-core reporting of cores not
released, fatal-before-completion publication, and a GM return gate released only
after close and its drain. Normal owners retire in parallel. Emergency retirement
collects all ready cores it wins into one broadcast and one deadline.

## Platform adaptation: fast-path window control

Not a cost decision. a5 has no fast-path enable anywhere in its software — no
offset, no `RegId` member, no `reg_offset` case, no AICore-side case, no
caller. Writing the dispatch register to idle is what opens a window there, and
there is no separately defined FAST_PATH close operation. This establishes the
software interface, not the absence of an equivalent on silicon; copying a2a3's
register offset without an a5 definition would be unjustified.

The earlier conclusion that this also made the GM return gate unnecessary was
too strong. The gate independently orders the worker's return after the AICPU's
last IDLE write, read-back and drain. TMR therefore uses isolated per-core gates,
reset before window-open. The worker acknowledges EXIT, then waits on its gate
before returning. Unacknowledged cores remain unreleased for host recovery.
The HBG resident and legacy paths use the same post-close gate rule. The
resident path keeps its cross-thread EXIT broadcast before each partition's
ACK sweep; the legacy path claims cores so normal and emergency retirement
cannot close or release the same worker twice.

## Historical claim-cost experiments

The claim is the whole cost. Taken as a2a3 takes it, per core, the port
measures **-12.35 us (0/27)**; everything else together measures
**-0.40 us (5/27)**, i.e. free — the sweep's batched broadcast pays for the
read-back. So three attempts went at the claim itself:

| Attempt | Result |
| ------- | ------ |
| Move the flags onto each core's own cache line (`CoreExecState` is already `alignas(64)` and the retiring thread has just read `reg_addr` from that line) | **Recovers half.** -12.35 to -7.03. The remaining cost is not false sharing |
| Relax the claim to `__ATOMIC_RELAXED` (sound: the loser only skips the core, and the winner's `reg_addr` was published at handshake) | **Nothing.** -8.56, no better than `acq_rel` |
| Take one claim per owning scheduler thread instead of one per core | **Works.** 8.24 us to 0.98 us |

The cost is thread scatter, not work, and it is proportional to how many times
the claim is taken — not to where the flags live or how they are ordered. A
probe arm with no claim at all isolates it: 8.24 us (27/27) between that arm
and the same build with a per-core claim, and the claim is the only thing that
inflates the broadcast phase, which contains no claim at all.

Disjoint ownership proves only that a per-owner claim can prevent duplicate
retirement. It does not prove equivalent emergency progress. The per-owner
implementation called the blocking platform helper once per owner, with a fresh
deadline each time. If the emergency caller alone handles several ready groups
with silent members, later groups receive EXIT only after earlier timeouts;
the wait can approach one budget per group. a2a3 broadcasts its entire claimed
emergency set first and spends one budget on that set. The port restores that
behavior and per-core arbitration, including overlapping target subsets.

The historical 8.24/0.98 us comparison does not price the current initialization
handoff or return gate, and is not a reason to weaken the fault protocol.

**Do not re-propose cache-line placement or memory ordering for this claim.**

## Dropped: two micro-optimizations of the close

| Attempt | Result |
| ------- | ------ |
| Split the close into a store burst then a read-back burst, so the read-backs pipeline instead of each draining its own store | **No distinguishable difference** (+0.05, 7/12). The predicted ~2 us from the STR/LDR round-trip cost table did not appear; that is not the bottleneck in this segment |
| Close each window on its ACK instead of deferring every close until the group has acknowledged | **No distinguishable difference** (+0.68, 21/27, below the segment's 1 us floor). The measured candidate kept the deferred close; the current A5 port also releases each acknowledged core's return gate only after the close/read-back/drain sequence |

## Ruled out on architecture, not measurement

- **A `dsb` instead of the read-back.** The window is Device-nGnRE; the E is
the early write-acknowledgement, which is what the barrier waits for.
- **One read-back covering the group.** nR orders accesses within one
peripheral, minimum 4 KB granule; each core has its own window.
- **Reading back `COND` rather than `DATA_MAIN_BASE`.** Same reason —
`DATA_MAIN_BASE` (0xD0) and `COND` (0x5108) are not in the same granule, so a
COND load does not order against the close.
- **A shared deadline with a blocking per-core wait** (what #2288 proposed).
Exit waiting has two independent properties: every core judged on its own
full budget, and the group bounded by one budget. The per-core wait has the
first, a shared deadline with a blocking wait has the second, and only the
non-blocking sweep has both.

## Where the earlier verdict went wrong

PR #2300 proposed recording this as "dropped, no net gain". Two of its findings
hold and are carried forward: the absence of a separately defined A5 FAST_PATH
register operation, and the
diagnostic round's eliminations (`runtime_destroy` unchanged, no extra COND
reads from the sweep). The verdict does not, for three reasons.

1. **"Only three steps have an a5 counterpart" undercounts.** That was measured
over the wait-and-close sequence alone. The claim, the read-back, its drain,
fatal-before-completion, and per-core reporting also port. The historical
candidate counted eight of ten, rather than three; the current port also
carries the independent GM return gate.
2. **The read-back was treated as tied to fast-path close and dropped with it.**
It is independent. a5's own close is a posted store to a Device-nGnRE window,
and the run publishes completion through a Normal-cacheable release store
that cannot order it.
3. **Its 59% transfer figure is specific to the per-core claim.** The earlier
per-thread candidate measured -1.43 us for the whole port; this does not
measure the current per-core claim, initialization handoff and return gate.
Comment thread
qweasdzxcht marked this conversation as resolved.

## Fault injection: where it had to move to

None of these mechanisms can be validated by a normal benchmark — their value
is entirely on the fault path, where a benchmark can only price them. The
first attempt drove the fault path from the host system tests and established
nothing: an unresponsive core forces a device reset, so the log line naming
the core never reaches the host, anything timed around the retirement is
dominated by the reset instead, and a test that counts retirement log lines
passes vacuously when no lines arrive at all. All three were withdrawn.

What replaced them is a device-side unit suite driving
`platform_retire_aicore_group` against simulated register blocks, in the style
the a2a3 suite already uses
(`tests/ut/cpp/a5/platform/test_aicore_retirement.cpp`).
Seven cases cover: the broadcast completing before any core is waited on; no
window closing until the group has acknowledged; a silent core left unclosed
while its answering peers are released; the `released[]` contract; one shared
budget for the group rather than one per core; address validation; and the
single-target group inheriting all of it. Three more cover the return
gates: release only after close and only to acknowledged cores, a gate held shut
while a peer's ACK is pending, and an invalid target rejecting the whole batch
before any EXIT is signalled. That is ten platform cases in total.

`tests/ut/cpp/a5/platform/test_return_gate_wait.cpp` adds three more: they include
the real a5sim `aicore/inner_kernel.h` and run the actual inline
`wait_for_post_close_release` on host threads. One waits on the gate word alone;
the other drives it against the real `platform_retire_aicore_group`, where the
worker stays held until the whole group has acknowledged and the production
retire path publishes the gate. The third covers HBG resident's autonomous
startup failure: with a stale previous-run RELEASE, it waits for the AICPU's
actual EXIT signal before ACK, then for close before return. All publish the
release before joining and
assert non-fatally, so a failed expectation cannot leave the un-timeouted wait
unjoined.

Each of the original seven was confirmed to discriminate by mutating the source
it tests — closing on acknowledgement, closing an unacknowledged core, retiring
serially with a per-core budget, reporting every core released, dropping the
validation, and bypassing the group from the then-existing single-core entry. Every mutant fails
the case that owns its property, and the unmutated source passes repeatedly. Two
of the seven were rewritten after the first mutation round showed they passed
against a serialized retirement and against that entry closing on
the signal alone. The three gate cases were added with the gate itself and are not
covered by that mutation round.

## Initialization and retirement ownership

An emergency can precede a peer's barrier-free assignment. Reading that
peer's tracker as a stable partition can retire an incomplete set, consume
its claim before its cores initialize, or retire the same cores again as
orphans. Clearing trackers between runs does not synchronize this run's
initialization.

Each core's atomic retirement state has READY and REQUESTED bits. Publishing

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Non-blocking follow-up: this READY/REQUESTED handoff adds a capability that A2A3's current single core_retired_ boolean does not have: preserving a retirement request that arrives before core initialization is published. Please track a separate PR to bring the corresponding A2A3 retirement paths to the same pending-retirement behavior, so the architectures do not retain different fatal-exit guarantees. This follow-up does not need to block the A5 work.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed this is a separate, non-blocking A2/A3 follow-up. I will first run the supplemental A2/A3 retirement experiment, then file a dedicated issue with the evidence and proposed acceptance criteria, and handle the A2/A3 change in its own PR. I am keeping that work out of the A5-focused #2388 commit.

READY releases initialization; requesting retirement acquires it.
The operation that observes only the other bit owns retirement. A request
before READY is serviced by the initializer, without requiring it to reach
the dispatch loop or normal shutdown. Barrier-free groups use the handshake's
fixed blocked partition even when tracker assignment fails. Serial startup
publishes assigned groups after its handshake barrier, or gives all cores to
one fallback publisher if assignment cannot complete. Emergency requests every
core before waiting, without reading a peer's partially assigned tracker. A
single retired boolean is insufficient here: skipping an unpublished core loses
a late request, while marking it retired consumes its claim before it is usable.
Pending requests are serviced after publication; no claim primitive guarantees
progress if the initializer itself never reaches publication.

`test_scheduler_retirement.cpp` in the A5 TMR unit directory drives the
production scheduler cold path with an observable retirement sink. Its twelve
cases cover late and partial initialization, concurrent publication and
retirement, normal/emergency exactly-once ownership, serial failure and success,
reset between generations, and fatal publication observed through completion.
The late and partial initialization cases fail against the original claim. It
also checks one emergency batch/deadline across owners, overlapping subsets, and
return-gate reset between generations, and
`ConcurrentOverlappingSetsClaimEachCoreOnce` has five threads retire arbitrary
overlapping, duplicated core sets from one start line and asserts every core is
retired exactly once — the per-core claim, not a per-caller one. Counts by file:
platform group 10, real sim inline wait 3, scheduler 12. Simulation checks
ordering, not physical MMIO completion.

## Still open

The read-back inside the window close has no observable effect on a simulated
register block; its justification remains the memory-attribute argument in
`docs/hardware/mmio-performance.md`. The scheduler tests use DFX disabled and
do not validate PMU finalization or device MMIO behavior on silicon.

On A5 DT device 0, the gated HBG candidate passed the selected empty,
mixed-chain, vector, and explicit legacy scene tests; the selected TMR
alternating and Dense16 cases also passed golden. The six focused C++ test
targets passed. This is normal-path coverage, not a hardware fault injection.
The official TMR benchmark ran each selected case for 50 rounds in one
main/candidate/main sequence on that device. Against the mean of the two main
arms, Effective changed by -10.79% for alternating Case1 and +2.76% for
Dense16. The latter is an unresolved workload-specific regression: the Sched
window includes AICore execution and dependency waits, so its change cannot
be attributed directly to the post-dispatch retirement tail.
Original file line number Diff line number Diff line change
Expand Up @@ -208,14 +208,24 @@ a5 TRB's `deinit` takes `Runtime * /*runtime*/` unnamed and performs none
| ---- | ------------------------------------ |
| a2a3 TRB | `dcci(my_hank, SINGLE_CACHE_LINE, CACHELINE_OUT)` (`:276`), then a **read** — bypass-load of `dev.teardown_gates[block_idx].post_close_release` (`:280`) |
| a2a3 HBG | same shape: `dcci` (`:301`), then teardown-gate read (`:305`) |
| a5 TRB | `dcci(my_hank, SINGLE_CACHE_LINE, CACHELINE_OUT)` (`:255`); no teardown gate exists on this variant |
| a5 TRB | at baseline `dcci(my_hank, SINGLE_CACHE_LINE, CACHELINE_OUT)` (`:255`), with no gate member on this variant; see the post-baseline note below |
| a5 HBG legacy | final `dcci(my_hank, SINGLE_CACHE_LINE, CACHELINE_OUT)` before return, commented "Flush all dirty cache lines to HBM before kernel exit" (`aicore_legacy_executor.cpp`, end of function) |
| a5 HBG resident | **no final handshake `dcci`**; the function ends at `write_reg(RegId::COND, AICORE_EXITED_VALUE)` (`aicore_executor.cpp:695`) |

> **Post-baseline change (#2388).** Everything here is stated at commit
> `31da0560e`, before the a2a3 retirement port. a5 TRB has since gained a
> `teardown_gates` member and now matches the a2a3 exit shape: `dcci(...,
> CACHELINE_OUT)` (`aicore_executor.cpp:268`) then a bypass load of
> `dev.teardown_gates[block_idx].post_close_release` (`:269`). Statements below
> that say a5 TRB "has no gate member" describe the baseline, not current code.
> The final #2388 revision also gates a5 HBG legacy and resident returns;
> resident startup/exit watchdog paths wait for the AICPU's actual EXIT before
> ACK so a reused descriptor's stale RELEASE cannot bypass the close.

The a2a3 paths' last descriptor access is a *read* placed after their own
write-back, which is a different exit shape from a5 TRB's write-back-and-return
and from a5 HBG resident's no-final-flush. The resident path's exit must not be
described using the legacy path's flush.
write-back. At this audit's baseline a5 TRB instead wrote back and returned, but
post-baseline (#2388) it takes the a2a3 shape; a5 HBG resident still has no final
flush. The resident path's exit must not be described using the legacy path's flush.

**What the primitive is** (`src/common/platform/onboard/aicpu/cache_ops.cpp:20`):

Expand All @@ -231,11 +241,12 @@ described using the legacy path's flush.
exercises any of this.

The AICPU **does store** into descriptor lines — `workers[i].task`
(`scheduler_cold_path.cpp`) and, on a2a3, `teardown_gates`. Whether any given line
is *dirty* at the `dc civac` call depends on the mapping's write policy and on
prior maintenance, neither of which is established here; so the clean half of that
call has a *possible* role in publishing AICPU-authored bytes, not a demonstrated
one. Either way, the call is not reducible to "invalidate for the next DMA".
(`scheduler_cold_path.cpp`) and, on a2a3 and post-baseline a5 TRB,
`teardown_gates`. Whether any given line is *dirty* at the `dc civac` call depends
on the mapping's write policy and on prior maintenance, neither of which is
established here; so the clean half of that call has a *possible* role in
publishing AICPU-authored bytes, not a demonstrated one. Either way, the call is
not reducible to "invalidate for the next DMA".

> **Withdrawn from the first version:** "the `dc civac` is performing a write-back
> of AICPU-authored bytes." Observed stores do not establish a dirty write-back
Expand Down Expand Up @@ -558,8 +569,12 @@ these:
cannot succeed at all, so the question there is recovery, not publication.
- **A core released but still spinning when the AICPU tears down** is the interleaving
above.
- **a5** has no gate member on TRB and does not use the gates on HBG, so this whole row
is a2a3-only.
- **a5** at this audit's baseline had no gate member on TRB and does not use the
gates on HBG, so this whole row is a2a3-only. Post-baseline (#2388) a5 TRB
carries per-core `teardown_gates`, and the final revision also uses them in
a5 HBG. The same failure and partial-exit questions apply; the ordering and
publication arguments above were written for a2a3 and are not restated here
as validated on a5.

**Limit of what is established about the failure path.** `retire_cores` logs each
unreleased core and returns `-1` (`scheduler_cold_path.cpp:695-703` TRB); this entry
Expand Down Expand Up @@ -603,7 +618,12 @@ in exactly this direction. It does **not** add up to "the clean has no consumer"
while the clean still spans `sizeof(runtime->dev)`, so the clean covers a device-only
region. **This is a post-`31da0560e` change**: at the older baseline
`runtime_device_copy_size` returned `sizeof(DeviceRuntimeLaunchDesc)` and the two
extents agreed. a5 TRB has no gate member, so its upload is the whole descriptor.
extents agreed. At baseline a5 TRB had no gate member, so its upload was the whole
descriptor. **Post-baseline (#2388)** a5 TRB gains `teardown_gates`, so its
`runtime_device_initialized_prefix_size` stops at
`offsetof(..., teardown_gates)` while `runtime_device_extent_size` stays
`sizeof(DeviceRuntimeLaunchDesc)` and the same two-extent split applies;
`runtime_device_copy_size` is unchanged.

### Evidence of no reader, versus unresolved reader

Expand Down
Loading
Loading