Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
78 commits
Select commit Hold shift + click to select a range
b541f19
feat(v41): add model metadata and text tokenizer entry
Sep 22, 2026
65b7472
feat(v41): add selective checkpoint loading and native weight packing
Sep 22, 2026
a0df435
feat(v41): prepare layer ownership and logical rank execution plans
Sep 22, 2026
39ddff0
feat(v41): define composite executor and runner lifecycle contracts
Sep 23, 2026
453461b
feat(v41): prepare bounded embeddings from TP vocabulary shards
Sep 23, 2026
399c729
feat(v41): track chunked prefill and request-owned cache state
Sep 23, 2026
deab8db
feat(v41): continue decode from committed prefill state
Sep 23, 2026
820ff9c
feat(v41): orchestrate rank-owned weights and complete layer composites
Sep 23, 2026
927e6aa
feat(v41): return validated logits to the shared generation loop
Sep 23, 2026
e09ce20
feat(v41): connect model loading and shared serving lifecycle
Sep 23, 2026
62c00b2
test(v41): supply explicit greedy sampling parameters
Sep 23, 2026
2a2e127
Add bounded V4.1 SWA and packed-FP4 MoE device segment
Sep 27, 2026
323bac6
Fix V4.1 cross-layer communication window lifecycle
Sep 27, 2026
a3b2d5f
Add real-weight SWA/MoE bundles and two-layer numerical validation
Sep 28, 2026
1633399
Use lib MoE heap budget in the real-weight diagnostic
Sep 28, 2026
b52934f
docs: register the V4.1 segment guide in navigation
Sep 28, 2026
2bfcdcf
Fix lib reference comparator invocation and allow saved-output checks
Sep 28, 2026
7ef97ca
Add post-run half-layer reference diagnostics without device feedback
Sep 28, 2026
0f59700
docs: record real-weight memory and numerical validation boundaries
Sep 28, 2026
c257c93
Keep stage and end-to-end numerical gates distinct and mandatory
Sep 28, 2026
d0a7e17
Add CPU trace replay to locate the first accumulated precision diverg…
Sep 28, 2026
3b06edc
Trace attention projection and quantization divergence on saved inputs
Sep 28, 2026
e33cc07
Add saved-boundary interventions for precision bisection
Sep 28, 2026
9d9ce9a
Allow selective internal tensor capture for precision diagnosis
Sep 28, 2026
5db8442
Support first-layer reference capture for matching device dumps
Sep 28, 2026
bc0023a
Preserve per-layer runtime dumps for dataflow bisection
Sep 28, 2026
466b5c0
Fix dump traversal before moving completed layer artifacts
Sep 28, 2026
83f7a17
Add independent attention reference precision sensitivity replay
Sep 28, 2026
80eb504
Validate SWA chain with checkpoint embedding input and replayable ini…
Sep 28, 2026
8d02950
test(v41): allow explicit token IDs in embedding precision control
Sep 28, 2026
6b4b873
fix(v41): compare decoded FP8 values in precision diagnostics
Sep 28, 2026
70191be
docs(v41): record text embedding precision control and remaining failure
Sep 28, 2026
85f3362
test(v41): adopt DSV4 layer residual budget and retain legacy profile
Sep 28, 2026
f74847b
test(v41): record approved accumulated pre-mix budget
Sep 28, 2026
f46d1f8
test(v41): probe FP8 projection reference accumulation sensitivity
Sep 28, 2026
7aa7918
test(v41): isolate attention RMSNorm reference reduction order
Sep 28, 2026
71e9929
feat(v41): pack fresh HC inputs into request-aware TP token slabs
Sep 28, 2026
b5f07cd
feat(v41): lower private SWA pages into causal chunk metadata
Sep 28, 2026
2ec6fd8
feat(v41): wire checkpoint RoPE and causal request pages into SWA val…
Sep 28, 2026
98a4311
fix(v41): replicate indexer heads required by current lib composites
Sep 28, 2026
78c0c52
feat(v41): dispatch prefill composites with independent communication…
Sep 28, 2026
32641e8
feat(v41): prepare producer-owned compressed attention weight bundles
Sep 28, 2026
b032b29
feat(v41): lower compressed pages and stable compressor state metadata
Sep 28, 2026
589c548
feat(v41): bind resident prefill producers across layer modes
Sep 28, 2026
6f52516
test(v41): validate real-weight C2A producer and reuse chain
Sep 28, 2026
fd0b762
test(v41): trace precision boundaries in each DP partition
Sep 28, 2026
9f9a5c3
refactor(v41): load attention probes without expert payloads
Sep 28, 2026
141c01a
Fix: retain V4.1 cache ownership across paused requests
Sep 28, 2026
85ee35e
Docs: record real-text precision candidate and page ownership limits
Sep 28, 2026
fabf011
Add: validate ragged C2A prefixes and empty DP groups
Sep 28, 2026
0355d40
Add: validate resident C2A compressor continuation across chunks
Sep 28, 2026
f28588d
docs: record unresolved projection precision experiment
Sep 28, 2026
861c36d
docs: record successful real-weight C2A continuation
Sep 28, 2026
4e83854
docs: record isolated V4.1 TP communication mismatch
Sep 29, 2026
043b7bf
docs(v41): record verified communication alignment fix
Sep 29, 2026
f3608ea
test(v41): prepare real-weight C1A producer chains
Sep 29, 2026
a11532c
fix(v41): keep C1A index argument write ranges disjoint
Sep 29, 2026
14e6a5a
test(v41): retain per-layer compressed attention dumps
Sep 29, 2026
8da7adc
docs(v41): record C1A device coverage and numerical failures
Sep 29, 2026
c5945eb
docs(v41): record C1A probability narrowing isolation
Sep 29, 2026
77e4844
docs(v41): correct C1A diagnostic reference and reject PV candidate
Sep 29, 2026
de14b45
docs(v41): record Q-A improvement and cache error propagation
Sep 29, 2026
9ad837e
test(v41): isolate C1A reuse attention replay
Sep 29, 2026
78530b8
docs(v41): record C1A continuation and six-layer results
Sep 29, 2026
f570f02
test(v41): preserve native gates for isolated kernel candidates
Sep 29, 2026
d19ba71
docs(v41): record isolated Q-A accumulation findings
Sep 29, 2026
9a60f59
docs(v41): close Q-A device and FP64 comparison
Sep 29, 2026
2c3c6f7
test(v41): cover bounded cross-page prefill continuation
Sep 29, 2026
4b39b38
test(v41): retain empty partitions in cross-page diagnostics
Sep 29, 2026
52cb5f9
docs(v41): record original Reuse Q-A control result
Sep 29, 2026
fac8eb1
docs(v41): record C2A cross-page device validation
Sep 29, 2026
0f3fb63
docs(v41): distinguish C1A crossing from validated C2A
Sep 29, 2026
c562231
docs(v41): record C1A cross-page device validation
Sep 29, 2026
5e7d564
Add V4.1 decode backbone and final output boundaries
Sep 29, 2026
a28ce2c
Map selective V4.1 weights into decode source slots
Sep 29, 2026
f59b6f3
Validate V4.1 decode weight geometry on checkpoint
Sep 29, 2026
959f959
Check every selected V4.1 decode weight ABI field
Sep 29, 2026
4c7b628
Validate V4.1 cache families before runner allocation
Sep 30, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
268 changes: 268 additions & 0 deletions docs/developer-guide/deepseek-v41-entry.md

Large diffs are not rendered by default.

532 changes: 532 additions & 0 deletions docs/developer-guide/v41-swa-segment.md

Large diffs are not rendered by default.

2 changes: 2 additions & 0 deletions mkdocs.yml
Original file line number Diff line number Diff line change
Expand Up @@ -111,3 +111,5 @@ nav:
- Weight Staging: developer-guide/weight-staging.md
- DeepSeek V4 Runtime: developer-guide/deepseek-v4-runtime.md
- DeepSeek V4 DSpark: developer-guide/deepseek-v4-dspark.md
- DeepSeek V4.1 Entry: developer-guide/deepseek-v41-entry.md
- DeepSeek V4.1 SWA Segment: developer-guide/v41-swa-segment.md
3 changes: 3 additions & 0 deletions pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -26,3 +26,6 @@ pypto-prepack-deepseek-v4 = "pypto_serving.tools.prepack_deepseek_v4:main"
where = ["."]
include = ["pypto_serving*"]
namespaces = false

[tool.setuptools.package-data]
"pypto_serving.model.deepseek_v41" = ["encoding.LICENSE"]
76 changes: 76 additions & 0 deletions pypto_serving/cli/main.py
Original file line number Diff line number Diff line change
Expand Up @@ -277,6 +277,8 @@ def build_serving_engine_config(args: argparse.Namespace) -> EngineConfig:
devices = parse_device_ids(args.devices, default_device=args.device)
model_config_data = read_model_config(model_dir)
model_family = detect_model_family(model_config_data)
if model_family == "deepseek_v41":
return _build_v41_engine_config(args, model_dir, model_config_data, devices)
model_variant = _resolve_model_variant(args)
_validate_prefill_chunk_size(
model_family,
Expand Down Expand Up @@ -350,6 +352,78 @@ def build_serving_engine_config(args: argparse.Namespace) -> EngineConfig:
)


def _build_v41_engine_config(args, model_dir, raw, devices):
"""Resolve the V4.1 worker boundary before starting any device resources.

Like DSpark placement, TP/DP are internal to one overlapped EP worker.
The scheduler uses two cache partitions instead of creating DP replicas.
"""
from pypto_serving.config.types import KVCacheGroupSpec
from pypto_serving.model.deepseek_v41.composite import (
MissingCompositeInterface, load_composite_bindings,
)
from pypto_serving.model.deepseek_v41.config import V41TextConfig
from pypto_serving.model.deepseek_v41.execution_plan import RankPlacement, plan_layers
from pypto_serving.serving.engine.async_engine import EngineConfig

text = V41TextConfig.from_dict(raw)
if getattr(args, "speculative_config", None) is not None or getattr(args, "num_speculative_tokens", 0):
raise ValueError("V4.1 text serving does not support DSpark or MTP")
if args.platform != "a5":
raise ValueError("V4.1 M0 serving requires --platform a5")
topology = (args.tensor_parallel_size, args.data_parallel_size, args.expert_parallel_size)
if topology != (4, 2, 8) or len(devices) != 8:
raise ValueError("V4.1 serving requires --tp 4 --dp 2 --ep 8 and exactly eight device IDs")
if args.block_size != 128:
raise ValueError("V4.1 serving requires --block-size 128")
if not 0 < args.max_model_len <= text.max_position_embeddings:
raise ValueError("V4.1 --max-model-len must fit the checkpoint position capacity")
if args.max_num_seqs <= 0 or args.max_num_batched_tokens <= 0:
raise ValueError("V4.1 batch and token capacities must be positive")

parallel = ParallelConfig(
data_parallel_size=1, tensor_parallel_size=1, expert_parallel_size=8,
enable_expert_parallel=True, placement_mode="overlapped", devices=devices,
data_parallel_routing=args.data_parallel_routing,
)
# This resolver intentionally raises while the lib adapter is missing.
# No default generic KV layout may be substituted for the V4.1 pools.
bindings = load_composite_bindings()
bindings.require(plan_layers(raw), RankPlacement(0))
groups = tuple(bindings.cache_groups)
if not groups or any(not isinstance(group, KVCacheGroupSpec) for group in groups):
raise MissingCompositeInterface("V4.1 bindings must provide concrete grouped cache specifications")
if any(group.num_partitions != 2 for group in groups):
raise ValueError("V4.1 cache groups must use two logical DP partitions")
if len({group.name for group in groups}) != len(groups):
raise ValueError("V4.1 cache group names must be unique")
runtime = dataclasses.replace(
_build_runtime_config(args),
kv_cache_groups=groups,
requires_homogeneous_prefill_decode=True,
)
return EngineConfig(
model_id=args.served_model_name or Path(model_dir).name,
model_dir=model_dir,
platform=args.platform,
device_id=devices[0],
device_ids=devices,
parallel_config=parallel,
executor_cls=_executor_cls_for_model_family("deepseek_v41"),
executor_kwargs={"use_compile_cache": args.use_compile_cache},
runtime_config=runtime,
profile_config=_build_profile_config(args),
max_num_running_reqs=args.max_num_seqs,
max_num_scheduled_tokens=args.max_num_batched_tokens,
long_prefill_token_threshold=args.long_prefill_token_threshold,
# Prefix restore needs both KV pages and matching compressor state.
# Pipelining needs independently owned mutable snapshots and tickets.
enable_prefix_cache=False,
async_scheduling=False,
enable_chunk_prefill=args.enable_chunked_prefill,
)


def _build_runtime_config(
args: argparse.Namespace,
*,
Expand Down Expand Up @@ -637,6 +711,8 @@ def _warn_deprecated_serving_profile_env(args: argparse.Namespace) -> None:

def _executor_cls_for_model_family(model_family: str, *, variant: str = "") -> str:
"""Map model family metadata to the worker executor class id."""
if model_family == "deepseek_v41":
return "PyptoDeepSeekV41Executor"
if model_family == "deepseek_v4":
if variant == "dspark":
return "PyptoDeepSeekV4DSparkExecutor"
Expand Down
3 changes: 3 additions & 0 deletions pypto_serving/config/types.py
Original file line number Diff line number Diff line change
Expand Up @@ -304,6 +304,9 @@ class PrefillBatch:
block_ids: list[list[int]] = field(default_factory=list)
block_ids_by_group: list[dict[str, list[int]]] = field(default_factory=list)
cache_partitions: list[int | None] = field(default_factory=list)
# Total original prompt lengths, distinct from seq_lens (this chunk's end).
# Required by integrations that own prefill-to-decode persistent state.
prompt_lens: list[int] = field(default_factory=list)


@dataclass
Expand Down
31 changes: 31 additions & 0 deletions pypto_serving/model/common/weights/store.py
Original file line number Diff line number Diff line change
Expand Up @@ -31,6 +31,10 @@ def get_tensor(self, name: str) -> torch.Tensor:
"""Return one tensor by name."""
raise NotImplementedError

def get_slice(self, name: str):
"""Return a lazy handle exposing shape/dtype and bounded slicing."""
raise NotImplementedError


class SafeOpenFn(Protocol):
"""Callable shape for injectable safetensors openers."""
Expand Down Expand Up @@ -110,6 +114,33 @@ def load_tensor(self, name: str) -> torch.Tensor:
"""Load one tensor by name, leaving all unrelated shard tensors untouched."""
return self.load_many([name])[name]

def load_slice(
self, name: str, ranges: tuple[slice, ...], *, shape: tuple[int, ...], dtype: str,
) -> torch.Tensor:
"""Validate metadata before reading an owned contiguous slice.

The caller supplies the expected physical storage shape and safetensors
dtype, and is responsible for its allocation budget. Existing whole-tensor
staging paths continue to use load_many.
"""
if len(ranges) != len(shape) or any(
not isinstance(part, slice) or part.step not in (None, 1)
or type(part.start) is not int or type(part.stop) is not int
or not 0 <= part.start < part.stop <= size
for part, size in zip(ranges, shape)
):
raise ValueError(f"invalid tensor slice: {name}")
path = self.path_for(name)
if not path.exists():
raise FileNotFoundError(self.missing_shard_error.format(path=path))
with self._safe_open_fn(path, self.device) as reader:
source = reader.get_slice(name)
if tuple(source.get_shape()) != shape or source.get_dtype() != dtype:
raise ValueError(f"checkpoint shape/dtype mismatch: {name}; expected {shape}/{dtype}")
# Clone even contiguous slices: do not retain mmap storage or a
# full source-row stride after the reader is closed.
return source[ranges].clone(memory_format=torch.contiguous_format)

def load_many(self, names: Sequence[str]) -> dict[str, torch.Tensor]:
"""Load a set of named tensors grouped by shard file.

Expand Down
10 changes: 10 additions & 0 deletions pypto_serving/model/deepseek_v41/__init__.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,10 @@
# Copyright (c) PyPTO Contributors.
# This program is free software, you can redistribute it and/or modify it under the terms and conditions of
# CANN Open Software License Agreement Version 2.0 (the "License").
# Please refer to the License for details. You may not use this file except in compliance with the License.
# THIS SOFTWARE IS PROVIDED ON AN "AS IS" BASIS, WITHOUT WARRANTIES OF ANY KIND, EITHER EXPRESS OR IMPLIED,
# INCLUDING BUT NOT LIMITED TO NON-INFRINGEMENT, MERCHANTABILITY, OR FITNESS FOR A PARTICULAR PURPOSE.
# See LICENSE in the root of the software repository for the full text of the License.
# -----------------------------------------------------------------------------------------------------------

"""DeepSeek V4.1 text model metadata and tokenization."""
69 changes: 69 additions & 0 deletions pypto_serving/model/deepseek_v41/cache_contract.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,69 @@
# Copyright (c) PyPTO Contributors.
# This program is free software, you can redistribute it and/or modify it under the terms and conditions of
# CANN Open Software License Agreement Version 2.0 (the "License").
# Please refer to the License for details. You may not use this file except in compliance with the License.
# THIS SOFTWARE IS PROVIDED ON AN "AS IS" BASIS, WITHOUT WARRANTIES OF ANY KIND, EITHER EXPRESS OR IMPLIED,
# INCLUDING BUT NOT LIMITED TO NON-INFRINGEMENT, MERCHANTABILITY, OR FITNESS FOR A PARTICULAR PURPOSE.
# See LICENSE in the root of the software repository for the full text of the License.
# -----------------------------------------------------------------------------------------------------------
"""Scheduler cache-family contract for the complete V4.1 text backbone.

The compressed pools have different source-token capacities per physical page:
C2A stores one row per two tokens, while C1A stores one row per token. Their
page IDs may be jointly lowered with index keys inside each family, but cannot
share one scheduler group with a single block-size declaration.
"""

from dataclasses import dataclass

from pypto_serving.config.types import KVCacheGroupSpec

from .execution_plan import LayerPlan


@dataclass(frozen=True)
class V41CacheGroups:
window: str
c2a: str
c1a: str


def validate_cache_groups(
layers: tuple[LayerPlan, ...], groups: tuple[KVCacheGroupSpec, ...], max_seq_len: int,
) -> V41CacheGroups:
"""Resolve the three full-history page families before device allocation."""
if len(layers) != 40 or tuple(layer.layer_id for layer in layers) != tuple(range(40)):
raise ValueError("V4.1 cache contract requires the complete ordered 40-layer backbone")
if type(max_seq_len) is not int or max_seq_len <= 0:
raise ValueError("V4.1 cache contract requires a positive sequence capacity")
if (len(groups) != 3 or any(not isinstance(group, KVCacheGroupSpec) for group in groups)
or len({group.name for group in groups}) != 3):
raise ValueError("V4.1 requires separate window, C2A and C1A cache groups")

expected = {
"window": (tuple(range(40)), 128, 1),
"c2a": (tuple(layer.layer_id for layer in layers
if layer.mode == "c2a_full" and layer.kv_source == layer.layer_id), 256, 2),
"c1a": (tuple(layer.layer_id for layer in layers
if layer.mode == "c1a_full" and layer.kv_source == layer.layer_id), 128, 1),
}
if not expected["c2a"][0] or not expected["c1a"][0]:
raise ValueError("V4.1 cache contract requires C2A and C1A KV producers")
resolved = {}
for group in groups:
matches = [family for family, (owners, _, _) in expected.items()
if tuple(group.layer_indices) == owners]
if len(matches) != 1 or matches[0] in resolved:
raise ValueError("V4.1 cache groups must match window and KV producer layers")
family = matches[0]
_, block_size, ratio = expected[family]
if (group.spec.block_size != block_size or group.spec.compress_ratio != ratio
or group.spec.storage_block_size != 128 or group.num_partitions != 2
or group.sliding_window is not None or group.is_eagle_group):
raise ValueError(f"V4.1 {family} cache page layout disagrees with the lib ABI")
required = (max_seq_len + block_size - 1) // block_size
if (group.max_blocks_per_seq < required
or (group.num_blocks is not None and group.num_blocks < required)):
raise ValueError(f"V4.1 {family} cache cannot cover max_seq_len")
resolved[family] = group.name
return V41CacheGroups(**resolved)
102 changes: 102 additions & 0 deletions pypto_serving/model/deepseek_v41/composite.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,102 @@
# Copyright (c) PyPTO Contributors.
# This program is free software, you can redistribute it and/or modify it under the terms and conditions of
# CANN Open Software License Agreement Version 2.0 (the "License").
# Please refer to the License for details. You may not use this file except in compliance with the License.
# THIS SOFTWARE IS PROVIDED ON AN "AS IS" BASIS, WITHOUT WARRANTIES OF ANY KIND, EITHER EXPRESS OR IMPLIED,
# INCLUDING BUT NOT LIMITED TO NON-INFRINGEMENT, MERCHANTABILITY, OR FITNESS FOR A PARTICULAR PURPOSE.
# See LICENSE in the root of the software repository for the full text of the License.
# -----------------------------------------------------------------------------------------------------------
"""Serving-owned composite boundary; no guessed lib signatures or CPU fallback.

An integration supplies complete layer entries (Attention plus FFN), resource
allocation and the model output boundary. Callbacks may enqueue work: wait()
returns only after all ranks have stopped accessing the supplied buffers.
"""
from dataclasses import dataclass
from typing import Callable, Mapping

from .execution_plan import LayerPlan, RankPlacement


class MissingCompositeInterface(NotImplementedError):
"""The selected lib revision cannot execute the requested serving segment."""


@dataclass(frozen=True)
class BuildOptions:
"""Worker-selected compilation settings passed unchanged to the adapter."""
platform: str = "a5"
pypto_build_dir: str = "build_output"
use_compile_cache: bool = False


@dataclass(frozen=True)
class LayerState:
"""Opaque device state carried between composites, never converted on host."""
residual: object
pre_mix: object
layout: str = "tp_replicated"


@dataclass(frozen=True)
class CompositeBindings:
"""Explicit adapter seam for a verified lib revision.

Entries consume (LayerPlan, LayerState, step, resources, weights) and return
LayerState. initialize consumes (embeddings, step, resources); output consumes
(state, step, resources) and returns host logits in original request order.
The adapter owns device uploads and lib ABI binding. Serving never expands
packed FP4 or calls the sub-operators of a complete layer.
allocate consumes (plan, device_ids, runtime, BuildOptions), including the
worker's build directory and compile-cache choice, and returns (resources,
num_pages). It also prepares global weights needed by initialize/output.

No default implementation fabricates results. A supplied adapter must handle
all DP partitions collectively, including empty partitions, and retain input
objects until wait() completes. reset_request must clear every persistent
cache and compressor slot for that request before returning.
"""
revision: str
entries: Mapping[tuple[str, str], Callable]
initialize: Callable
output: Callable
allocate: Callable
prepare_weights: Callable
reset_request: Callable
wait: Callable
close: Callable
cache_groups: tuple = ()
input_layout: str = "tp_replicated"
output_layout: str = "tp_replicated"
decode_backbone: Callable | None = None

def require(self, layers: tuple[LayerPlan, ...], placement: RankPlacement) -> None:
if not self.revision:
raise ValueError("composite bindings must identify the validated lib revision")
phases = ("prefill",) if callable(self.decode_backbone) else ("prefill", "decode")
missing = sorted({f"{phase}/{layer.mode}" for phase in phases
for layer in layers if not callable(self.entries.get((phase, layer.mode)))})
if missing:
raise MissingCompositeInterface("missing complete layer composites: " + ", ".join(missing))
for name in ("initialize", "output", "allocate", "prepare_weights", "reset_request", "wait", "close"):
if not callable(getattr(self, name)):
raise MissingCompositeInterface(f"missing composite resource/output operation: {name}")
if self.input_layout not in ("tp_replicated", "tp_local_token") or self.output_layout != self.input_layout:
raise MissingCompositeInterface("layer entries must preserve their declared residual/pre_mix layout")
if (placement.tp_size, placement.dp_size, placement.ep_size) != (4, 2, 8):
raise ValueError("V4.1 serving currently targets TP4/DP2/EP8")


def load_composite_bindings() -> CompositeBindings:
"""Fail before resource allocation until a complete adapter is implemented.

This is an intentional integration placeholder, not a discovery heuristic:
importing a lib module or finding a function does not establish its ABI.
"""
raise MissingCompositeInterface(
"V4.1 serving execution requires verified lib composite bindings: "
"all-mode prefill/decode adapters, initial residual/pre_mix, "
"cache allocation/reset/completion and final HC/Norm/LM head. "
"The bounded SWA Attention/MoE segment is available separately; it is not a complete model backend. "
"Track pypto-lib #1205, #1275 and #1287; no Torch fallback is enabled."
)
Loading
Loading