Skip to content

runtime: a cublaslt bf16 gemm launch may pin its algorithm - #25

Merged
xiaguan merged 2 commits into
pegainfer-project:masterfrom
FeathBow:runtime/wide-gemm
Sep 30, 2026
Merged

xiaguan merged 2 commits into
pegainfer-project:masterfrom
FeathBow:runtime/wide-gemm

Conversation

@FeathBow

@FeathBow FeathBow commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

What

Reworked as the review asked: the runtime executes the algorithm the manifest names and decides nothing at run time. extern:cublaslt_bf16_tn and extern:cublaslt_bf16_tn_acc take an optional algo, every cublasLtMatmulAlgoConfigAttributes_t value of one algorithm: id, tile, stages, split_k, reduction, swizzle, custom, inner_shape, cluster_shape. Without algo nothing changes.

  • Load. cublasLtMatmulAlgoInit and every attribute set, then cublasLtMatmulAlgoCheck for each call of the launch at the shapes its args take with every var at its lower bound and at its upper one, the when var held to the launch's range, with the workspace within Blas's 32 MiB. An algorithm that cannot be built or does not fit is a kernel artifact error, never a fallback.
  • Issue. cublasLtMatmul on the pinned algorithm with Blas's workspace: no heuristic, no sync, no readback, so eager multi-rank issue, kern bench captures and when-skipped launches need nothing from it.
  • Per shape. One launch per when range, each with its own algo.
  • The verifier refuses algo on any other extern. docs/runtime.md records how a caller chooses one and the sm_89 finding that SPLITK_NUM 1 does not make two algorithms agree bitwise (the sliced kernels, algo 30, 31 and 16). Schema regenerated.

Verification

  • tests/gemm_algo.rs (GPU, ignored like fp8_gemm.rs): an op with two when ranges of rows, each pinned to cuBLASLt's first heuristic answer at the end of its range, lands the default op's exact product (a permutation times small integers) at 100 and 256 rows, eagerly and through a captured graph; a tile the shape cannot run fails the load as a kernel artifact error.
  • gemm_algo and fp8_gemm pass on Single GPU (sm_89, x86_64) and Single GPU (GH200, aarch64), both CUDA 13.1.
  • cargo fmt --check, cargo clippy --all-targets -- -D warnings, cargo test -p kern-manifest -p kern-pool -p kern-runtime -p kern-test and the schema golden pass.
  • End to end in pega-omni's HiDream-O1 engine (hidream-o1: serve hidream-o1-image-dev as a kern manifest on one gpu pega-omni#6), which pins its four decoder GEMMs from a file its tuner measures on the card: its golden test passes on the pinned algorithms, and through its server a 2048 x 2048 picture takes 15.56 and 15.55 s pinned against 19.65 and 19.67 s on the heuristic on the sm_89 card, 4.24 and 4.25 s against 4.24 and 4.24 s on the GH200.

@FeathBow FeathBow changed the title runtime: a bf16 gemm built-in on cublaslt's 256x128 tile runtime: a tuned bf16 gemm built-in over one bitwise class of cublaslt algorithms Sep 25, 2026
@xiaguan

xiaguan commented Sep 26, 2026

Copy link
Copy Markdown
Contributor

Thanks for this, and for the measurements. The sliced-kernel finding on sm_89 (algo 30/31/16 report SPLITK_NUM 1 yet differ bitwise) is exactly the kind of thing we want recorded.

We'd like to land the capability in a different shape, though. The runtime should not measure or decide at run time. It should only execute what the manifest says. Planning inside the runtime breaks things in practice:

  • kern bench captures through profile.rs (Probe::program, and now Probe::without for --ablate) without going through plan_tuned. An unplanned shape there fails with first met inside a graph capture.
  • In an eager multi-rank enqueue, the plan's synchronize() inside launch waits on an in-flight collective whose peer hasn't been issued yet, and deadlocks.
  • Since rsi: harness fixes, write-up and site for the vLLM optimization loop #29, plan_tuned also plans launches that when would skip. It also conflicts with master in exec.rs.

We don't need to switch algorithms online, so the proposal is to pin the algorithm in the manifest:

  • extern:cublaslt_bf16_tn (and _acc) takes an optional algo config: id, tile, stages, split-K, reduction scheme, swizzle, custom option, cluster shape, i.e. the cublasLtMatmulAlgoConfigAttributes.
  • At load: cublasLtMatmulAlgoInit + set the attributes + cublasLtMatmulAlgoCheck against the launch's shape. A mismatch is an artifact error, never a silent fallback. At issue: no heuristic, no sync, no readback. With no algo, behavior is unchanged.
  • Per-shape choice uses when from rsi: harness fixes, write-up and site for the vLLM optimization loop #29: one extern launch per m-range, each with its own algo.
  • A GPU test in the style of fp8_gemm.rs: a pinned algo runs and lands the default op's exact product.

How the algo is chosen stays outside kern for now. You can keep timing candidates in pega-omni's real workload and write the winner into the manifest. With the algorithm pinned, kern test against the reference gates the output as usual. Happy to review a revised PR along these lines.

Signed-off-by: Feathbow <feathbow@gmail.com>
@FeathBow FeathBow changed the title runtime: a tuned bf16 gemm built-in over one bitwise class of cublaslt algorithms runtime: a cublaslt bf16 gemm launch may pin its algorithm Sep 26, 2026
@FeathBow

Copy link
Copy Markdown
Contributor Author

Thanks, reworked along these lines. The runtime now only executes what the manifest pins: cublaslt_bf16_tn[_acc] takes an optional algo, built and AlgoChecked at load over the launch's when range (a mismatch is a kernel artifact error), with no heuristic, sync or readback at issue. The online tuning is gone. pega-omni#6 measures the algorithms per card outside kern and writes them into the manifest: on sm_89 a picture takes 15.6 s pinned against 19.7 s on the heuristic, and on a GH200 the two are the same. GPU test passes on both.

Signed-off-by: Feathbow <feathbow@gmail.com>
@xiaguan
xiaguan merged commit b43bb5a into pegainfer-project:master Sep 30, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants