Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
87 changes: 87 additions & 0 deletions docs/user-guide/deepseek-v4-dspark-tool-constraints.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,87 @@
# DeepSeek V4 DSpark constrained tool deployment

This guide deploys structural tool-call constraints for DeepSeek V4 Flash DSpark with fused K7 decoding. The tool names and JSON schemas come from each `/v1/chat/completions` request, not from a server-side allowlist or a particular agent client. Serving compiles the request grammar on the host; the device receives a fixed-layout allowed-token mask and still runs one fused K7 decode dispatch per step.

## Compatible components

- Use a DeepSeek V4 Flash DSpark W8A8 checkpoint and the 16-device `--dp 4 --ep 16 --tp 4` topology described in [DSpark serving](../developer-guide/deepseek-v4-dspark.md).
- Deploy PyPTO Serving and PyPTO-Lib as a matched pair. Their prefill and decode positional ABI now includes `grammar_mask`; fused K7 decode also includes `valid_draft_counts`. Updating only one repository will not work. The mask shape and argument order are fixed regardless of whether a request has constraints.
- The mask also carries a row-mode header in segment 0's last padding word: `-1` selects ordinary argmax and `0` enables masking. Serving packs this header; it is not part of the provider's vocabulary bitset. Deploy the matching packer and sampler together even when their tensor shapes have not changed.
- Install `xgrammar==0.2.7` into the Python environment that runs Serving. This is the version validated with the DeepSeek V4 structural-tag grammar and the checkpoint tokenizer. XGrammar is a host-side dependency; PyPTO-Lib does not import it. See the [XGrammar package](https://pypi.org/project/xgrammar/0.2.7/).
- DSpark currently supports greedy sampling only. Use `temperature: 0`; non-greedy requests are rejected rather than silently changing sampling behavior.

```bash
python -m pip install 'xgrammar==0.2.7'
python -m pip show xgrammar
```

When upgrading either repository, use a new `PYPTO_PROG_BUILD_DIR` for the changed kernels before enabling `--use-compile-cache`. The compile cache does not validate a previous executable against the new source or positional ABI.

## Start the service

The paths below are deployment choices, not required repository locations. Set the checkpoint and build-cache paths for your environment. The 16 listed devices must be free before starting the server.

```bash
PYPTO_PROG_BUILD_DIR=/path/to/new-compile-cache \
python -m pypto_serving.cli \
--model /path/to/dsv4-flash-dspark-w8a8 \
--served-model-name dsv4-flash-dspark-w8a8 \
--backend npu --platform a2a3 \
--devices 0,1,2,3,4,5,6,7,8,9,10,11,12,13,14,15 \
--dp 4 --ep 16 --tp 4 --block-size 32 \
--max-model-len 16384 --max-num-seqs 8 \
--max-num-batched-tokens 8192 \
--long-prefill-token-threshold 128 \
--speculative-config '{"method":"dspark","num_speculative_tokens":7}' \
--enable-prefix-caching \
--ring-heap 2147483648,2147483648,4294967296,8589934592 \
--generate-config '{"max_new_tokens":2048,"temperature":0}' \
--use-compile-cache --host 127.0.0.1 --port 8000
```

Start without `--use-compile-cache` for the first build if your deployment does not use a persistent compile directory. Check `/health` before sending requests:

```bash
curl --noproxy '*' http://127.0.0.1:8000/health
```

## Verify constrained and ordinary requests

This request requires a `shell` tool call with a `command` string and forbids extra keys. The client still decides whether and how to execute the returned command.

```bash
curl --noproxy '*' http://127.0.0.1:8000/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "dsv4-flash-dspark-w8a8",
"messages": [{"role":"user","content":"Show the current directory."}],
"tools": [{"type":"function","function":{
"name":"shell",
"description":"Run a shell command",
"strict":true,
"parameters":{
"type":"object",
"properties":{"command":{"type":"string"}},
"required":["command"],
"additionalProperties":false
}
}}],
"tool_choice":"required",
"temperature":0,
"max_tokens":256
}'
```

Expect a completed `message.tool_calls` entry named `shell`, with `function.arguments` as a JSON string containing `command`, and `finish_reason: "tool_calls"`. A length-truncated call is not a successful invocation. For a baseline, repeat the request with `tool_choice: "auto"` and `strict` omitted: it follows the unconstrained path and may answer directly or produce a tool call. `/v1/completions` uses the same fixed kernel ABI with an all-allowed mask.

`auto` enters structural constraints only when at least one declared tool has `strict: true`; `required` and a named tool choice also enter the constrained path. A schema needs `required` to require a field and `additionalProperties: false` to exclude unknown fields. Ordinary non-strict `auto` does not guarantee schema-valid parameters. `parallel_tool_calls` is passed to the grammar only for constrained requests.

Ordinary batches reuse immutable device-resident all-allowed masks and default draft counts. The sampler skips bitset decoding for ordinary rows; constrained rows apply the mask before the same running-maximum reduction. These paths share the fused K7 interface and do not require a separate forward/sampling dispatch.

Constrained requests reuse request-local Host bitmask storage. Draft validation and mask generation share one speculative traversal, then roll back; only committed output advances the grammar. Serving packs the valid prefix and bonus row directly into shared buffers without unpacking vocabulary bits. With asynchronous scheduling enabled, acceptance-independent metadata can prepare early, but mask planning still waits for that request's previous output to be committed. Mutable slot ownership lasts through output reclaim. Constrained masks still use the existing full-tensor runtime upload path; Host row reuse is not partial H2D transfer.

Missing XGrammar, an unsupported schema, or an unsupported model path yields an HTTP 400 before streaming headers are sent. There is no fallback to unconstrained generation for a request that asked for constraints. Request completion and cancellation release request-local grammar state. Recompute preemption of a constrained request is not yet supported; under cache pressure it may wait for resources rather than preempt another constrained request.

## Upgrade and rollback

Roll Serving and PyPTO-Lib forward or backward together, then use a build-cache directory compiled from that pair. After a restart, check health, one constrained tool request, and one ordinary chat request before restoring traffic. Do not reuse a compile cache across the changed kernel ABI or mix the new Serving `TaskArgs` order with old Lib kernels.
8 changes: 4 additions & 4 deletions docs/user-guide/online-serving.md
Original file line number Diff line number Diff line change
Expand Up @@ -58,7 +58,7 @@ The server converts chat messages to a prompt with the tokenizer's `apply_chat_t

## DeepSeek V4 Function Tools

Tool calling is selected by the model tokenizer; no extra launcher flag or vLLM dependency is required. Serving encodes tool definitions and parses model output. **The client executes tools**, then sends the results in a new chat request.
Tool calling is selected by the model tokenizer; no extra launcher flag or vLLM dependency is required. Serving encodes tool definitions and parses model output. **The client executes tools**, then sends the results in a new chat request. Constrained generation for DeepSeek V4 DSpark K7 additionally requires XGrammar; see [deployment and compatibility](deepseek-v4-dspark-tool-constraints.md).

Send this request to a **DeepSeek V4** server, not the Qwen server in the examples above. Set `DEEPSEEK_BASE_URL` to that server's host and port (8000 is the default serving port).

Expand All @@ -84,7 +84,7 @@ curl --noproxy "*" "$DEEPSEEK_BASE_URL/v1/chat/completions" \

`auto` is the default when non-empty `tools` are supplied. The model can answer normally or return `message.tool_calls`, with each call containing `id`, `type: "function"`, and `function: {name, arguments}`. `arguments` is a **JSON string**, not a JSON object. `content` can be null; `reasoning`, when enabled, remains separate from both content and tools.

The parser returns a model-generated function name even if that name is absent from this request's `tools`, matching vLLM's default DeepSeek V4 behavior. Serving does not provide built-in functions such as `read_file`; the client decides which calls it can execute. Check the returned name against the client's available tools before executing it.
In ordinary unconstrained `auto` mode, the parser returns a model-generated function name even if that name is absent from this request's `tools`, matching vLLM's default DeepSeek V4 behavior. Serving does not provide built-in functions such as `read_file`; the client decides which calls it can execute. Check the returned name against the client's available tools before executing it.

For a successful call, the client should validate the function name and arguments against its schema before execution. Append the returned assistant message and a tool result that references the same call ID:

Expand All @@ -109,8 +109,8 @@ Send that history to the same chat endpoint, including `tools` again if another
Supported controls and limits:

- `tool_choice: "none"` suppresses tool-call output. Recognized tool blocks are consumed when tools are supplied; ordinary no-tools chat keeps its existing parser behavior.
- Multiple calls are supported. `parallel_tool_calls: false` exposes only the first call; it does not constrain sampling or execute tools serially.
- `required`, named tool choices, and `strict: true` return HTTP 400 because constrained tool decoding is not implemented. Non-function tool types are rejected during request validation.
- Multiple calls are supported. In unconstrained `auto`, `parallel_tool_calls: false` exposes only the first parsed call; in constrained DSpark K7 generation, the flag also enters the structural grammar. It never executes tools on the server.
- DeepSeek V4 DSpark K7 supports structural constraints for `required`, a named tool choice, and `auto` when at least one tool has `strict: true`. These modes require XGrammar and the matching PyPTO-Lib kernel ABI. Ordinary `auto` with no strict tool remains unconstrained. Unsupported model paths or unavailable XGrammar reject constrained requests before generation. Non-function tool types are rejected during request validation.
- Tools on a model without a registered tool parser are rejected. DSML formatting stays in the DeepSeek implementation, not the HTTP server or scheduler.
- Tool-history argument values cannot contain the reserved `</|DSML|parameter>` delimiter, including inside nested JSON values. Tool-result content cannot contain `</tool_result>`. These inputs return HTTP 400 before generation rather than breaking the history encoding.
- The parser preserves DSML parameter types without schema-based coercion or guessed JSON repairs. A length-truncated call can have incomplete arguments: do not execute it as a successful call.
Expand Down
1 change: 1 addition & 0 deletions mkdocs.yml
Original file line number Diff line number Diff line change
Expand Up @@ -94,6 +94,7 @@ nav:
- Inference and Serving:
- Offline Inference: user-guide/offline-inference.md
- Online Serving: user-guide/online-serving.md
- DSpark Tool Constraints: user-guide/deepseek-v4-dspark-tool-constraints.md
- Parallelism and Scaling: user-guide/parallel.md
- Benchmarking: user-guide/benchmarking.md
- Profiling: user-guide/profile.md
Expand Down
2 changes: 1 addition & 1 deletion pypto-lib
Submodule pypto-lib updated 68 files
+10 −56 .github/actions/setup-ci-job/action.yml
+16 −5 .github/scripts/run_a5_pytest.py
+3 −10 .github/workflows/ci.yml
+2 −3 .github/workflows/daily_ci.yml
+99 −12 docs/models/deepseek_v4_1_flash/index.md
+2 −3 models/deepseek_v4_1_flash/_golden_smoke.py
+47 −0 models/deepseek_v4_1_flash/attention_ops.py
+184 −0 models/deepseek_v4_1_flash/attention_sp.py
+10 −3 models/deepseek_v4_1_flash/attention_tp.py
+282 −0 models/deepseek_v4_1_flash/boundary_fwd.py
+10 −1 models/deepseek_v4_1_flash/config.py
+2 −2 models/deepseek_v4_1_flash/decode_attn_c1a_full.py
+196 −169 models/deepseek_v4_1_flash/decode_attn_c2a_full.py
+70 −62 models/deepseek_v4_1_flash/decode_attn_c2a_reuse.py
+1 −1 models/deepseek_v4_1_flash/decode_c1a_full.py
+1 −1 models/deepseek_v4_1_flash/decode_c1a_reindex.py
+1 −1 models/deepseek_v4_1_flash/decode_c1a_reuse.py
+8 −13 models/deepseek_v4_1_flash/decode_c2a_full.py
+8 −13 models/deepseek_v4_1_flash/decode_c2a_reuse.py
+4 −3 models/deepseek_v4_1_flash/decode_common.py
+4,302 −0 models/deepseek_v4_1_flash/decode_fwd.py
+1,063 −86 models/deepseek_v4_1_flash/decode_layer.py
+0 −104 models/deepseek_v4_1_flash/decode_layer_plan.py
+8 −13 models/deepseek_v4_1_flash/decode_swa.py
+338 −192 models/deepseek_v4_1_flash/ep_transport.py
+121 −143 models/deepseek_v4_1_flash/expert_routed.py
+138 −0 models/deepseek_v4_1_flash/hc_head.py
+10 −7 models/deepseek_v4_1_flash/hc_mixes.py
+40 −5 models/deepseek_v4_1_flash/hc_post.py
+3 −3 models/deepseek_v4_1_flash/hc_pre.py
+177 −0 models/deepseek_v4_1_flash/input_pack.py
+845 −0 models/deepseek_v4_1_flash/lm_head.py
+123 −71 models/deepseek_v4_1_flash/moe.py
+38 −1 models/deepseek_v4_1_flash/o_proj.py
+97 −98 models/deepseek_v4_1_flash/prefill_attn_c1a_full.py
+50 −51 models/deepseek_v4_1_flash/prefill_attn_c1a_reindex.py
+82 −75 models/deepseek_v4_1_flash/prefill_attn_c1a_reuse.py
+299 −13 models/deepseek_v4_1_flash/prefill_attn_c2a_full.py
+59 −11 models/deepseek_v4_1_flash/prefill_attn_c2a_reuse.py
+46 −9 models/deepseek_v4_1_flash/prefill_attn_swa.py
+2 −2 models/deepseek_v4_1_flash/prefill_c1a_common.py
+1 −1 models/deepseek_v4_1_flash/prefill_c1a_full.py
+15 −15 models/deepseek_v4_1_flash/prefill_c1a_indexer.py
+1 −1 models/deepseek_v4_1_flash/prefill_c1a_reindex.py
+1 −1 models/deepseek_v4_1_flash/prefill_c1a_reuse.py
+854 −0 models/deepseek_v4_1_flash/prefill_c1a_sp.py
+322 −84 models/deepseek_v4_1_flash/prefill_c2a_full.py
+113 −22 models/deepseek_v4_1_flash/prefill_c2a_reuse.py
+45 −18 models/deepseek_v4_1_flash/prefill_layer.py
+263 −97 models/deepseek_v4_1_flash/prefill_swa.py
+2 −2 models/deepseek_v4_1_flash/rmsnorm.py
+5 −5 models/deepseek_v4_flash_dspark/decode_cp_allgather.py
+2 −4 models/deepseek_v4_flash_dspark/decode_csa.py
+10 −2 models/deepseek_v4_flash_dspark/decode_fwd.py
+13 −5 models/deepseek_v4_flash_dspark/decode_fwd_dspark.py
+17 −17 models/deepseek_v4_flash_dspark/decode_prepare.py
+1 −2 models/deepseek_v4_flash_dspark/decode_swa.py
+66 −49 models/deepseek_v4_flash_dspark/dspark_attention.py
+129 −41 models/deepseek_v4_flash_dspark/lm_head.py
+10 −2 models/deepseek_v4_flash_dspark/prefill_fwd.py
+2 −1 models/deepseek_v4_flash_dspark/prefill_hca.py
+4 −4 models/deepseek_v4_flash_dspark/prefill_sparse_attn.py
+6 −8 models/deepseek_v4_flash_dspark/qkv_proj_rope.py
+18 −9 models/deepseek_v4_flash_dspark/rmsnorm.py
+92 −0 tests/ci/test_run_a5_pytest.py
+482 −62 tests/contract/test_deepseek_v4_1_flash_contract.py
+207 −0 tests/contract/test_v41_prefill_c2a_sp.py
+148 −0 tests/contract/test_v41_prefill_swa_sp.py
3 changes: 3 additions & 0 deletions pypto_serving/config/types.py
Original file line number Diff line number Diff line change
Expand Up @@ -304,6 +304,8 @@ class PrefillBatch:
block_ids: list[list[int]] = field(default_factory=list)
block_ids_by_group: list[dict[str, list[int]]] = field(default_factory=list)
cache_partitions: list[int | None] = field(default_factory=list)
# Opaque request-local generation constraints, interpreted only by capable runners.
constraint_states: dict[str, object] = field(default_factory=dict)


@dataclass
Expand Down Expand Up @@ -335,6 +337,7 @@ class DecodeBatch:
block_ids: list[list[int]] = field(default_factory=list)
block_ids_by_group: list[dict[str, list[int]]] = field(default_factory=list)
cache_partitions: list[int | None] = field(default_factory=list)
constraint_states: dict[str, object] = field(default_factory=dict)
# Optional MTP context for models (e.g. DeepSeek V4) that decode two real
# trailing tokens per step. ``prev_token_ids`` holds the token id at absolute
# position ``seq_len-2`` per request (shape ``[B]``) and ``prev_hidden_states``
Expand Down
Loading
Loading