[None][fix] Scope KVCM warmup capacity constraints to DeepSeek V4 - #19213
yizhang-nv wants to merge 15 commits into
Conversation
|
/bot run --disable-fail-fast --extra-stage "DGX_B200-PyTorch-Post-Merge-1,DGX_B200-PyTorch-Post-Merge-2,DGX_H100-PyTorch-Post-Merge-1,DGX_H100-PyTorch-Post-Merge-2" |
|
PR_Github #73564 [ run ] triggered by Bot. Commit: |
|
/bot run --disable-fail-fast --stage-list "A10-PyTorch-1,A10-PyTorch-2,A10-PyTorch-3,DGX_H100-PyTorch-1,DGX_H100-PyTorch-2,DGX_H100-PyTorch-3,DGX_H100-PyTorch-4,DGX_H100-PyTorch-5,DGX_H100-PyTorch-6,DGX_B200-PyTorch-Post-Merge-1,DGX_B200-PyTorch-Post-Merge-2,DGX_H100-PyTorch-Post-Merge-1,DGX_H100-PyTorch-Post-Merge-2" |
|
PR_Github #73573 [ run ] triggered by Bot. Commit: |
|
PR_Github #73564 [ run ] completed with state |
|
PR_Github #73573 [ run ] completed with state
|
61d98aa to
72adcea
Compare
|
/bot run --disable-fail-fast |
|
PR_Github #73797 [ run ] triggered by Bot. Commit: |
This comment has been minimized.
This comment has been minimized.
This comment has been minimized.
This comment has been minimized.
Semantic conflict reviewThe verdict of record is the Latest recorded state: No semantic conflict found (best effort) for head Best-effort AI judgment for the recorded revisions. PASS, FAIL and INCONCLUSIVE may be incomplete or incorrect. PR authors and reviewers should independently verify the evidence and relevant behavior. This semantic review and its status/workflow are advisory, not required merge checks under current repository rules; other merge requirements still apply. Advisory status does not make a confirmed defect safe to ignore.
Processed request and reply comments are minimized to reduce timeline noise; they remain expandable for audit. |
This comment has been minimized.
This comment has been minimized.
This comment has been minimized.
This comment has been minimized.
This comment has been minimized.
This comment has been minimized.
This comment has been minimized.
This comment has been minimized.
Signed-off-by: Yi Zhang <187001205+yizhang-nv@users.noreply.github.com>
Signed-off-by: Yi Zhang <187001205+yizhang-nv@users.noreply.github.com>
Signed-off-by: Yi Zhang <187001205+yizhang-nv@users.noreply.github.com>
Signed-off-by: Yi Zhang <187001205+yizhang-nv@users.noreply.github.com>
Signed-off-by: Yi Zhang <187001205+yizhang-nv@users.noreply.github.com>
Signed-off-by: Yi Zhang <187001205+yizhang-nv@users.noreply.github.com>
Signed-off-by: Yi Zhang <187001205+yizhang-nv@users.noreply.github.com>
Signed-off-by: Yi Zhang <187001205+yizhang-nv@users.noreply.github.com>
…ek V4 Signed-off-by: Yi Zhang <187001205+yizhang-nv@users.noreply.github.com>
Signed-off-by: Yi Zhang <187001205+yizhang-nv@users.noreply.github.com>
Signed-off-by: Yi Zhang <187001205+yizhang-nv@users.noreply.github.com>
Signed-off-by: Yi Zhang <187001205+yizhang-nv@users.noreply.github.com>
Signed-off-by: Yi Zhang <187001205+yizhang-nv@users.noreply.github.com>
Restore CUDA graph batch clipping when estimating V4's long-decode cache constraint. Keep the generation dummy input counted once and align the allocation tests with that contract, including the actual graph builder. Validate dynamic draft schedules against observed batch sizes while requiring coverage of each configured draft length, without depending on exact request admission thresholds. Validation: 154 cases passed on a B200 using the CI 62214 wheel with the V4 runtime fix overlaid, including 15 attention graph capture/replay cases. Signed-off-by: Yi Zhang <187001205+yizhang-nv@users.noreply.github.com>
Signed-off-by: Yi Zhang <187001205+yizhang-nv@users.noreply.github.com>
abb6079 to
45fc4ac
Compare
|
/bot run --disable-fail-fast |
|
PR_Github #76185 [ run ] triggered by Bot. Commit: |
|
PR_Github #76185 [ run ] completed with state
|
|
/bot run --disable-fail-fast |
|
PR_Github #76191 [ run ] triggered by Bot. Commit: |
|
PR_Github #76191 [ run ] completed with state |
Description
DeepSeek V4's maximum-sequence-length warmup constraint was moved into generic KVCM V2 in #16545. For ordinary attention models, that hard floor can enlarge the temporary KV pool beyond its estimated GPU budget, causing OOM during cache creation or encoder profiling.
Restore the longest-decode-plus-short-requests constraint to
DeepseekV4CacheManager. Keep the general context/chunked-prefill constraint in generic V2: it covers the configured per-iteration token budget. Preserve V4's draft/extra reservations, explicit pool-ratio opt-out, and generic average-length pool preferences.Also correct generation dummy allocation at the source.
token_numsalready includes history plus the current input; the generation resize counted that input again. Remove only this duplicate+1, retaining draft/extra reservations and the normal scheduler's generation growth. This lets CUDA-graph warmup use the capacity returned by the manager without skipping feasible batch shapes.model_engine.pyis unchanged from main.Original gRPC, Seed-OSS, Mistral, and multimodal-example fixtures explicitly select V2. A benchmark comment is corrected without changing its budget. No native allocator or public configuration changes.
Test Coverage
The committed generic test changes retain the context constraint assertions and add one regression with draft lengths 0 and 4: query the available capacity, then allocate a generation dummy exactly at a page boundary. Both cases fail on the old dummy allocator and pass with the correction. The enclosing executor directory is already included in
l0_h100.yml.One CPU-only V4 configuration test covers the default ratio and an explicit pool ratio. It asserts the exact longest-decode-plus-minimal-decodes constraint (including draft/extra reservations), preserves the inherited context constraint, and checks the explicit-ratio opt-out. The existing
l0_cpu.ymlattention-directory entry collects itscpu_onlymarker. Diagnostic scripts and larger temporary matrices remain outside the PR.Real B200 validation on September 22 PDT / September 23 UTC:
StorageStatistics.The Seed pytest completed successfully and workers shut down cleanly; its outer temporary shell runner subsequently exited 2 because that script was edited while the long test ran. This harness error and its correction are preserved in the evidence report; it is separate from the passing pytest/JUnit result.
These are isolated Python-policy comparisons on the matching CI60862 native/Python runtime (
40466ac6c0), using original model tests frozen ateb7be9db6a; they are not a fresh native build of the rebased branch (63e64e5bdb). Earlier same-runtime validation reproduced and fixed the original A10 gRPC, B200 Seed, and H100 Mistral OOMs. A10/H100 full-model tests were not repeated for this final dummy-token correction; dynamic-tree and Helix consumers were checked in source, not rerun on hardware.Evidence, module hashes, full stdout/stderr, and JUnit paths:
tmp/v4-relocation/dummy-token-fix/RESULTS.mdandresults-summary.jsonin the PR workspace; full logs under/home/scratch.yizhan_sw_1/logs/2026-09-22/.PR Checklist
GitHub Bot Help
To see available CI bot commands, comment
/bot help.