feat(v41): stage text loader and prefill composite integration - #258
hashiqiqixian wants to merge 78 commits into
Conversation
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. Important Review skippedAuto incremental reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the ⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Advanced Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Advanced Run ID: 📒 Files selected for processing (16)
Included review availability: Your plan provides up to 1 included review per hour; 0 remain after this review. 📝 WalkthroughWalkthroughThe pull request adds DeepSeek V4.1 metadata detection, text configuration validation, prompt encoding, local tokenizer support, packaging, documentation, and tests. Model loading and serving execution remain explicitly unsupported. ChangesDeepSeek V4.1 support
Priority: ⬇️ Low Estimated code review effort: 3 (Moderate) | ~25 minutes Change: Feature Sequence Diagram(s)sequenceDiagram
participant Requester
participant ModelFamily
participant ModelLoader
participant DeepSeekV41TokenizerAdapter
participant encode_messages
Requester->>ModelFamily: detect_model_family(config)
ModelFamily-->>Requester: deepseek_v41
Requester->>ModelLoader: load(model_dir)
ModelLoader->>ModelLoader: load_text_config(model_dir)
ModelLoader-->>Requester: NotImplementedError
Requester->>DeepSeekV41TokenizerAdapter: apply_chat_template(messages)
DeepSeekV41TokenizerAdapter->>encode_messages: encode validated messages
encode_messages-->>DeepSeekV41TokenizerAdapter: prompt string
Merge Risk: ⚪ Minimal · up to This change adds staged V4.1 metadata and tokenizer support while clearly rejecting unsupported serving and weight-loading paths, so no actionable merge-blocking risk remains. 🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
Full details: Docstring CoverageExplanation Docstring coverage is 22.86% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 35 functions across 10 files. (6 skipped: 6 unsupported.) Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. A rabbit checks the V4.1 guide, Comment |
Keep the saved SWA state immutable while selecting causal request prefixes. Pass matching global and local active counts through metadata, device dispatch and native references, including odd compressor tails and empty partitions.
Prepare causal source slices and metadata for two consecutive request chunks while retaining device cache and compressor handles. Compare each native stage after device execution and keep captured outputs out of the forward input path.
Problem and resulting behavior
The upstream serving stack did not recognize DeepSeek V4.1 Flash or have a
checkpoint-to-lib data path. This PR registers its text configuration and
tokenizer, selectively reads the real checkpoint, and adds a V4-style
Loader → Executor → Runner framework with explicit lib-composite bindings.
Unsupported model execution fails before device allocation; there is no Torch
fallback or small-operator production chain.
The bounded prefill adapter invokes existing lib Attention half-layer and
packed-FP4 MoE composites while carrying FP32 residual/pre_mix on device. It
prepares rank-local weights, initial embeddings, causal positions, RoPE rows,
private physical window/compressed/index pages and compressor state. Producer
bindings connect SWA, C2A Full/Reuse and C1A Full/Reindex/Reuse. Each Attention
entry retains its own communication epoch; MoE advances across layers.
Request ownership and commit/reset contracts are represented by a ledger.
Engram, vision, DSpark/MTP and tool calling are excluded. Their weights are
not loaded. The default full-model backend remains disabled until all mode,
decode, output and resource contracts are verified.
Validation
empty-partition cross-page tests and all pre-commit checks passed at the
latest commit. Real checkpoint CPU checks covered 93,366 text tensor
headers, 54 selected payload/conversion cases and bounded 8K embeddings.
a 31+1 continuation with empty DP passed 48/48. C1A Full20/Reuse21 31+1
continuation passed 60/60. These use resident caches and windows.
end-to-end HC and Reuse22 Attention still fail. The failing Reuse row was
reproduced in an isolated replay with the unchanged reference; original
Q-A and compensated Q-A both fail the same row. FP64 captures identify
reference/device accumulation-rounding amplification, but the original
gate remains failed.
per-rank V4 residual gate. Accumulated pre_mix still has 3/256 elements
outside the approved provisional rtol=0.01, atol=0.001 gate. A compensated
Q-A experiment left all saved tensors unchanged. No precision candidate
or tolerance change is promoted by this PR.
--repeat-input-chunksnow provides bounded, explicit cross-page statediagnostics. Host tests cover causal reads, disjoint writes and an empty
DP group. A5 C2A single-request 9×32 continuation now passes 216/216
native checks over 36 stages, including the second compressed page;
the inactive DP cache remains untouched. C1A single-request 5×32 crossing also completed with 150/150
native checks over 20 stages. Its second window/KV/index pages
were populated and the inactive DP cache stayed untouched.
This used the Full-only Q-A diagnostic lib revision and does not
override prior two-active-DP C1A precision failures. The repeated boundary
input is injected, not a complete preceding model segment.
Device comparisons use the specialized C1A native reference (BF16 PV and
its first-vector patch), not a generic FP32 sparse-attention reference.
Diagnostic cuts and FP64 checks do not replace acceptance references. Detailed
revisions, metrics and rejected experiments are in
docs/developer-guide/v41-swa-segment.md.The original official Reuse Q-A control also fails the same row 31;
all its actual and expected tensors are bitwise equal to the
compensated candidate. Reverting that Q-A change does not pass the
original native gate.
Remaining boundaries
transition through the device resource adapter.
and emitting final HC+Norm.
boundary_embed_to_normrepacks embeddings andcannot consume the completed backbone state.
validate real-checkpoint 8192 prefill → 128 decode on TP4/DP2/EP8. The
current A5 account whitelist permits only cards 0–3, so four-card results
do not establish M0.
This PR does not change the lib or compiler pin. The compiler's 64-byte
physical communication-buffer fix is tracked in pypto#2942 and validated only
on the matched diagnostic stack; it is not merged here. Serving issue #240
tracks the overall integration.