Summary
Establish Simpler's program execution path as a safe, measurable backend for vLLM serving. The work starts from the Pipeline A run-to-run enqueue capability and ends with a fixed Qwen workload running through a vLLM eager adapter, followed by capture/replay support where the target SDK and platform contracts are proven.
The device remains serial at the operator level. The host may prepare and enqueue later runs while an earlier run is still executing; every run must retain independent completion, result, error, and resource ownership until its last consumer is finished.
Why this issue exists
Pipeline A and the P3 contracts provide the runtime foundation for this work:
This issue is the end-to-end integration tracker. Individual implementation PRs should remain focused and link back here.
Scope and boundaries
The work is split into a program-path milestone and a later vLLM/capture milestone.
Program-path milestone
- Validate the current P4 capability: real
Worker.submit early enqueue, whole-operator device ordering, sustained refill, and depth-one default behavior.
- Revalidate P1–P3 contracts under the new admission path:
- run-owned parameter, runtime/image, arena, result, timing, and diagnostic resources;
- per-run completion/drain, error attribution, event references, and generation retirement;
- parent-run dispatch FIFO, prepared-successor admission, build/recorder ownership, and host accessor ordering;
- safe default, unsupported-shape fallback, diagnostic exclusivity, direct-control, close, and stop-admission paths.
- Extend the proven path only where the target workload requires it, with an explicit capability matrix for A3/A5 HBG, TMR, HOST/DEVICE tensors, group/SUB/communication/provider shapes, and endpoint scope.
- Replace fixed-slot assumptions with bounded resource accounting and backpressure only after run identity, last-consumer retirement, workspace growth, and failure handling are proven.
- Integrate continuous DFX collection by collector and runtime, preserving run identity, buffer ownership, accounting, bounded flush/close, and diagnostic isolation.
vLLM milestone
- Freeze the Qwen workload and ABI: weights, prompt/token hash, prefill/decode shape, KV/page layout, sampling, output contract, and device/software versions.
- Add a fixed-shape Simpler-backed vLLM eager path through the public model-runner boundary, then cover the agreed dynamic batch and request-lifecycle range.
- Add capture/replay only after the preparation-result, parameter-binding, graph-reference lifetime, and target SDK contracts are established. Capture/replay is a separate capability boundary from the first program-path milestone.
- Run end-to-end correctness and performance comparisons with identical workload and device conditions.
Current acceptance gate
The first P4 acceptance phase passed on the bounded A3 HBG scope: local L3, same endpoint, HOST tensors, single non-group NEXT_LEVEL task, launch_depth=2, with depth one as the default control.
The P1–P3 regression phase is currently partial:
- CPU contract tests passed, including admission/lifecycle coverage after the current runtime was built.
- A3
worker_async_fifo passed for HBG and TMR.
worker_async_endpoint failed on both device 0 and device 1 with the same FFTSPLUS AICore fault (507018, 0x4000000000000000), followed by scheduler -100 and force reset.
The original endpoint failure was investigated with controlled depth-one/depth-two runs and five-round pressure tests on two additional devices. The fault did not reproduce with the same main, callable, and runtime. It is retained as a non-stable device/runtime anomaly; no code repair is opened from this incident. The bounded program-path milestone is now complete for the documented scope, while vendor-level root-cause analysis remains optional follow-up if maintainers request it.
Milestones
Step 2 closeout and Step 3 handoff
The bounded Qwen Step 2 scope is complete in two review PRs:
- #2447: real Qwen depth-one
Worker.submit, 127-step autoregressive token/KV qualification, adapter and lifecycle contract.
- #2456: workload manifest, dependency matrix, depth-two request evidence, and explicit safe serial fallback when host sampled-token feedback prevents joined enqueue.
Both PRs are merged into main. The Step 2 depth-two conclusion is a safe fallback; it does not claim joined native early enqueue.
Step 3 starts with the L3 clone-request path on the current program runtime. The first proof uses multiple logically independent requests with identical workload content:
- request content may be identical;
- request/run/generation identities must differ;
- live KV pages, block tables, input staging, sampled-token/output buffers, events, errors, results, and resource leases must be request-private;
- request B must be prepared and enqueued while request A is still executing;
- operator execution remains serial, parent dispatch FIFO remains ordered, and A remains readable after B starts;
- physical capacity may be reused only after the previous request's last consumer retires.
PR #2471 is the first Step 3 contract/evidence PR. It adds the clone-request identity/resource checks to the existing A3 early-enqueue path and records A3 hardware evidence. Its current head is 0095497523c33fb86766d5e90e01de649ea578a3, based on main at 833327f60ea8f4772bf1673ea8edb349314a350d. The complete GitHub CI matrix passed, including A2A3/A5 simulation and onboard checks, UT, pre-commit, build, and DeepSeek/network validation. The PR remains open pending merge. It does not add Qwen adapter code or change runtime admission.
After the clone-request contract is accepted, the project will connect the real Qwen adapter on the post-merge Step 2 baseline. Same-request adjacent decode feedback (sampled token/KV/sequence metadata) is a separate DEVICE producer/consumer capability and remains depth-one fallback until proven. A5 HBG, TMR, multi-worker extensions, and distinct-request HBG numerical correctness remain separate capability work. The existing program-path boundaries remain in force: operator execution stays serial, host access has no implicit cross-run synchronization, and unsupported shapes must fall back or be rejected explicitly.
Tracking rules
Every implementation PR must:
- link this issue in the PR body, using
Part of #<issue-number> unless the PR is intended to close the entire issue;
- state the exact phase and capability it changes;
- record the base/head, supported scope, tests, hardware task IDs and evidence paths;
- preserve the distinction between accepted, prepared, enqueued, completed, and retired states;
- document unsupported or fallback paths when the change affects admission.
Progress updates on this issue should summarize the phase, current commit, evidence, open blocker, and next action. Individual PR merge does not close this issue. Close this issue only after the milestone checklist and final vLLM acceptance are complete.
Related design boundaries
The first milestone follows the existing program-path Pipeline A design. Kernel API work, tiling caches, ACLGraph, and capture-specific resource contracts are separate follow-on capabilities and are not implicit prerequisites for the bounded program-path acceptance. No stage in this issue changes device execution from serial operator ordering to device concurrency, adds implicit host tensor synchronization, or promises a performance gain before controlled measurements are complete.
Summary
Establish Simpler's program execution path as a safe, measurable backend for vLLM serving. The work starts from the Pipeline A run-to-run enqueue capability and ends with a fixed Qwen workload running through a vLLM eager adapter, followed by capture/replay support where the target SDK and platform contracts are proven.
The device remains serial at the operator level. The host may prepare and enqueue later runs while an earlier run is still executing; every run must retain independent completion, result, error, and resource ownership until its last consumer is finished.
Why this issue exists
Pipeline A and the P3 contracts provide the runtime foundation for this work:
5818054b) is merged and provides the first boundedWorker.submitearly-enqueue capability.This issue is the end-to-end integration tracker. Individual implementation PRs should remain focused and link back here.
Scope and boundaries
The work is split into a program-path milestone and a later vLLM/capture milestone.
Program-path milestone
Worker.submitearly enqueue, whole-operator device ordering, sustained refill, and depth-one default behavior.vLLM milestone
Current acceptance gate
The first P4 acceptance phase passed on the bounded A3 HBG scope: local L3, same endpoint, HOST tensors, single non-group
NEXT_LEVELtask,launch_depth=2, with depth one as the default control.The P1–P3 regression phase is currently partial:
worker_async_fifopassed for HBG and TMR.worker_async_endpointfailed on both device 0 and device 1 with the same FFTSPLUS AICore fault (507018,0x4000000000000000), followed by scheduler-100and force reset.The original endpoint failure was investigated with controlled depth-one/depth-two runs and five-round pressure tests on two additional devices. The fault did not reproduce with the same main, callable, and runtime. It is retained as a non-stable device/runtime anomaly; no code repair is opened from this incident. The bounded program-path milestone is now complete for the documented scope, while vendor-level root-cause analysis remains optional follow-up if maintainers request it.
Milestones
Worker.submitpath is correct for the bounded Step 2 scope (depth one real qualification; depth-two safe fallback)Step 2 closeout and Step 3 handoff
The bounded Qwen Step 2 scope is complete in two review PRs:
Worker.submit, 127-step autoregressive token/KV qualification, adapter and lifecycle contract.Both PRs are merged into
main. The Step 2 depth-two conclusion is a safe fallback; it does not claim joined native early enqueue.Step 3 starts with the L3 clone-request path on the current program runtime. The first proof uses multiple logically independent requests with identical workload content:
PR #2471 is the first Step 3 contract/evidence PR. It adds the clone-request identity/resource checks to the existing A3 early-enqueue path and records A3 hardware evidence. Its current head is
0095497523c33fb86766d5e90e01de649ea578a3, based onmainat833327f60ea8f4772bf1673ea8edb349314a350d. The complete GitHub CI matrix passed, including A2A3/A5 simulation and onboard checks, UT, pre-commit, build, and DeepSeek/network validation. The PR remains open pending merge. It does not add Qwen adapter code or change runtime admission.After the clone-request contract is accepted, the project will connect the real Qwen adapter on the post-merge Step 2 baseline. Same-request adjacent decode feedback (sampled token/KV/sequence metadata) is a separate DEVICE producer/consumer capability and remains depth-one fallback until proven. A5 HBG, TMR, multi-worker extensions, and distinct-request HBG numerical correctness remain separate capability work. The existing program-path boundaries remain in force: operator execution stays serial, host access has no implicit cross-run synchronization, and unsupported shapes must fall back or be rejected explicitly.
Tracking rules
Every implementation PR must:
Part of #<issue-number>unless the PR is intended to close the entire issue;Progress updates on this issue should summarize the phase, current commit, evidence, open blocker, and next action. Individual PR merge does not close this issue. Close this issue only after the milestone checklist and final vLLM acceptance are complete.
Related design boundaries
The first milestone follows the existing program-path Pipeline A design. Kernel API work, tiling caches, ACLGraph, and capture-specific resource contracts are separate follow-on capabilities and are not implicit prerequisites for the bounded program-path acceptance. No stage in this issue changes device execution from serial operator ordering to device concurrency, adds implicit host tensor synchronization, or promises a performance gain before controlled measurements are complete.