Skip to content

feat(v41): stage text loader and prefill composite integration - #258

Open
hashiqiqixian wants to merge 78 commits into
hw-native-sys:mainfrom
hashiqiqixian:feat/v41-model-entry
Open

hashiqiqixian wants to merge 78 commits into
hw-native-sys:mainfrom
hashiqiqixian:feat/v41-model-entry

Conversation

@hashiqiqixian

@hashiqiqixian hashiqiqixian commented Sep 22, 2026 •

Copy link
Copy Markdown

Problem and resulting behavior

The upstream serving stack did not recognize DeepSeek V4.1 Flash or have a
checkpoint-to-lib data path. This PR registers its text configuration and
tokenizer, selectively reads the real checkpoint, and adds a V4-style
Loader → Executor → Runner framework with explicit lib-composite bindings.
Unsupported model execution fails before device allocation; there is no Torch
fallback or small-operator production chain.

The bounded prefill adapter invokes existing lib Attention half-layer and
packed-FP4 MoE composites while carrying FP32 residual/pre_mix on device. It
prepares rank-local weights, initial embeddings, causal positions, RoPE rows,
private physical window/compressed/index pages and compressor state. Producer
bindings connect SWA, C2A Full/Reuse and C1A Full/Reindex/Reuse. Each Attention
entry retains its own communication epoch; MoE advances across layers.
Request ownership and commit/reset contracts are represented by a ledger.

Engram, vision, DSpark/MTP and tool calling are excluded. Their weights are
not loaded. The default full-model backend remains disabled until all mode,
decode, output and resource contracts are verified.

Validation

  • Local: 305 V4.1 unit tests passed; the added
    empty-partition cross-page tests and all pre-commit checks passed at the
    latest commit. Real checkpoint CPU checks covered 93,366 text tensor
    headers, 54 selected payload/conversion cases and bounded 8K embeddings.
  • A5 TP2/DP2/EP4: C2A layers 2–3 passed 24/24 native stage/state checks;
    a 31+1 continuation with empty DP passed 48/48. C1A Full20/Reuse21 31+1
    continuation passed 60/60. These use resident caches and windows.
  • C1A layers 20–25 with MoE executed and passed 88/90 native checks. Full20
    end-to-end HC and Reuse22 Attention still fail. The failing Reuse row was
    reproduced in an isolated replay with the unchanged reference; original
    Q-A and compensated Q-A both fail the same row. FP64 captures identify
    reference/device accumulation-rounding amplification, but the original
    gate remains failed.
  • A two-layer real-text SWA chain passes its 18 native half-layer checks and
    per-rank V4 residual gate. Accumulated pre_mix still has 3/256 elements
    outside the approved provisional rtol=0.01, atol=0.001 gate. A compensated
    Q-A experiment left all saved tensors unchanged. No precision candidate
    or tolerance change is promoted by this PR.
  • --repeat-input-chunks now provides bounded, explicit cross-page state
    diagnostics. Host tests cover causal reads, disjoint writes and an empty
    DP group. A5 C2A single-request 9×32 continuation now passes 216/216
    native checks over 36 stages, including the second compressed page;
    the inactive DP cache remains untouched. C1A single-request 5×32 crossing also completed with 150/150
    native checks over 20 stages. Its second window/KV/index pages
    were populated and the inactive DP cache stayed untouched.
    This used the Full-only Q-A diagnostic lib revision and does not
    override prior two-active-DP C1A precision failures. The repeated boundary
    input is injected, not a complete preceding model segment.

Device comparisons use the specialized C1A native reference (BF16 PV and
its first-vector patch), not a generic FP32 sparse-attention reference.
Diagnostic cuts and FP64 checks do not replace acceptance references. Detailed
revisions, metrics and rejected experiments are in
docs/developer-guide/v41-swa-segment.md.

The original official Reuse Q-A control also fails the same row 31;
all its actual and expected tensors are bitwise equal to the
compensated candidate. Reverting that Q-A change does not pass the
original native gate.

Remaining boundaries

  • Resolve accumulated precision without changing the approved comparator.
  • Verify every compressed mode, producer, continuation, reset/reuse and decode
    transition through the device resource adapter.
  • Obtain a lib composite accepting the existing final residual/pre_mix
    and emitting final HC+Norm. boundary_embed_to_norm repacks embeddings and
    cannot consume the completed backbone state.
  • Connect final head/greedy feedback and the shared HTTP lifecycle, then
    validate real-checkpoint 8192 prefill → 128 decode on TP4/DP2/EP8. The
    current A5 account whitelist permits only cards 0–3, so four-card results
    do not establish M0.

This PR does not change the lib or compiler pin. The compiler's 64-byte
physical communication-buffer fix is tracked in pypto#2942 and validated only
on the matched diagnostic stack; it is not merged here. Serving issue #240
tracks the overall integration.

@coderabbitai

coderabbitai Bot commented Sep 22, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

Important

Review skipped

Auto incremental reviews are disabled on this repository.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Advanced

Run ID: 17fabe55-9324-45a9-81c3-8aca6760251b

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Advanced

Run ID: 9b3cfaba-8110-468d-9c0d-271f0cf17ffc

📥 Commits

Reviewing files that changed from the base of the PR and between 306a776 and b541f19.

📒 Files selected for processing (16)
  • docs/developer-guide/deepseek-v41-entry.md
  • mkdocs.yml
  • pyproject.toml
  • pypto_serving/cli/main.py
  • pypto_serving/model/deepseek_v41/__init__.py
  • pypto_serving/model/deepseek_v41/config.py
  • pypto_serving/model/deepseek_v41/encoding.LICENSE
  • pypto_serving/model/deepseek_v41/encoding.py
  • pypto_serving/model/deepseek_v41/tokenizer.py
  • pypto_serving/model/model_family.py
  • pypto_serving/model/model_loader.py
  • pypto_serving/model/tokenizer.py
  • tests/fixtures/deepseek_v41/chat_encoding_golden.json
  • tests/fixtures/deepseek_v41/config.json
  • tests/unit/model/deepseek_v41/test_encoding.py
  • tests/unit/model/deepseek_v41/test_entry.py

Included review availability: Your plan provides up to 1 included review per hour; 0 remain after this review.


📝 Walkthrough

Walkthrough

The pull request adds DeepSeek V4.1 metadata detection, text configuration validation, prompt encoding, local tokenizer support, packaging, documentation, and tests. Model loading and serving execution remain explicitly unsupported.

Changes

DeepSeek V4.1 support

Layer / File(s) Summary
Model detection and text configuration
pypto_serving/model/model_family.py, pypto_serving/model/deepseek_v41/config.py, tests/fixtures/deepseek_v41/config.json
Detects DeepSeek V4.1 metadata and validates text-backbone fields before creating an immutable V41TextConfig.
Prompt encoding and tokenizer adapter
pypto_serving/model/deepseek_v41/encoding.py, pypto_serving/model/deepseek_v41/tokenizer.py, tests/fixtures/deepseek_v41/chat_encoding_golden.json, tests/unit/model/deepseek_v41/test_encoding.py
Adds validated chat-message rendering, reasoning-effort handling, thinking-mode behavior, and local fast-tokenizer integration.
Runtime dispatch and execution rejection
pypto_serving/model/model_loader.py, pypto_serving/model/tokenizer.py, pypto_serving/cli/main.py
Routes DeepSeek V4.1 checkpoints to text configuration loading and the dedicated tokenizer adapter. Loader and CLI paths raise NotImplementedError before weight loading or serving execution.
Validation, packaging, and developer guide
tests/unit/model/deepseek_v41/test_entry.py, pyproject.toml, docs/developer-guide/deepseek-v41-entry.md, mkdocs.yml
Adds entry-level coverage, packages encoding.LICENSE, and documents the staged integration scope. Imports and license files are added for the new module.

Priority: ⬇️ Low

Estimated code review effort: 3 (Moderate) | ~25 minutes

Change: Feature

Sequence Diagram(s)

sequenceDiagram
  participant Requester
  participant ModelFamily
  participant ModelLoader
  participant DeepSeekV41TokenizerAdapter
  participant encode_messages
  Requester->>ModelFamily: detect_model_family(config)
  ModelFamily-->>Requester: deepseek_v41
  Requester->>ModelLoader: load(model_dir)
  ModelLoader->>ModelLoader: load_text_config(model_dir)
  ModelLoader-->>Requester: NotImplementedError
  Requester->>DeepSeekV41TokenizerAdapter: apply_chat_template(messages)
  DeepSeekV41TokenizerAdapter->>encode_messages: encode validated messages
  encode_messages-->>DeepSeekV41TokenizerAdapter: prompt string
Loading

Merge Risk: ⚪ Minimal · up to b541f

This change adds staged V4.1 metadata and tokenizer support while clearly rejecting unsupported serving and weight-loading paths, so no actionable merge-blocking risk remains.

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 22.86% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 35 functions across 10 files. (6 skipped:… Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description check ✅ Passed The description covers DeepSeek V4.1 configuration, tokenizer support, execution gating, and validation. It also includes prefill integration details that are not represented in the summarized changes…
Linked Issues check ✅ Passed The pull request identifies serving issue #240 as related and correctly states that the issue is not closed by this change.
Out of Scope Changes check ✅ Passed The summarized changes stay within the stated scope. They do not add weight loading, inference execution, experimental backends, native attention adapters, Engram support, or library pin changes.
Title check ✅ Passed The title refers to real V4.1 text-loader staging, but its prefill composite integration wording is not supported by the summarized changes. It is partially related rather than fully descriptive.
Full details: Docstring Coverage

Explanation

Docstring coverage is 22.86% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 35 functions across 10 files. (6 skipped: 6 unsupported.)


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

A rabbit checks the V4.1 guide,
Config fields stand firm inside,
Prompts hop through tokens bright,
Tests guard each encoded byte,
Execution waits for future light.

Comment @coderabbitai help to get the list of available commands.

@hashiqiqixian hashiqiqixian changed the title feat(v41): add model metadata and text tokenizer entry feat(v41): add model entry and selective checkpoint loading Sep 22, 2026
@hashiqiqixian hashiqiqixian changed the title feat(v41): add model entry and selective checkpoint loading feat(v41): add model entry, selective weights and execution planning Sep 22, 2026
@hashiqiqixian hashiqiqixian changed the title feat(v41): add model entry, selective weights and execution planning feat(v41): scaffold staged serving integration with composite contracts Sep 23, 2026
@hashiqiqixian hashiqiqixian changed the title feat(v41): scaffold staged serving integration with composite contracts feat(v41): add staged serving framework and SWA/MoE device segment Sep 27, 2026
ChenShenAi added 21 commits September 29, 2026 03:33
Keep the saved SWA state immutable while selecting causal request prefixes. Pass matching global and local active counts through metadata, device dispatch and native references, including odd compressor tails and empty partitions.
Prepare causal source slices and metadata for two consecutive request chunks while retaining device cache and compressor handles. Compare each native stage after device execution and keep captured outputs out of the forward input path.
@hashiqiqixian hashiqiqixian changed the title feat(v41): add staged serving framework and SWA/MoE device segment feat(v41): stage text loader and prefill composite integration Sep 29, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant