[None][fix] Bound V2 disaggregated KV transfer admission - #19855
Draft
yizhang-nv wants to merge 2 commits into
Draft
yizhang-nv wants to merge 2 commits into
yizhang-nv wants to merge 2 commits into
Conversation
Signed-off-by: Yi Zhang <187001205+yizhang-nv@users.noreply.github.com>
Signed-off-by: Yi Zhang <187001205+yizhang-nv@users.noreply.github.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
The asynchronous Python transfer path with KVCM V2 and PP=1 bypasses the executor transfer window. At high disaggregated concurrency, inline receive setup repeatedly waits for bounce-buffer reservations on the GEN execution thread, starving decode launches. GPT-OSS 120B FP4 at ISL8192/OSL1024 and concurrency1024 reproduced a 21.92% output-throughput regression against a measured V1 baseline.
Restore the transfer window and share its budget with the V2 scheduler before KV allocation. Defer new receives without starving active decode, preserve admission order among KV-fitting requests and oversized-request progress, and distinguish transfer backpressure from KV exhaustion. The early path is limited to asynchronous Python V2/PP1; other configurations keep post-scheduling admission. Applying only a late window caused repeated V2 allocation/rollback, which the early budget avoids.
Related: #17495. No public configuration/API, dependency, workload, test selection, or waiver changes.
Test Coverage
Validated on two ARM GB200 nodes with CTX TP1 / GEN TP4 from main
fc0876cfd6c5d661f186707857a4e5bfeb12bc6f, using a clean native/bindings/wheel build and verified runtime imports. The Python patch retains identical native artifacts and dependencies. Eight uninstrumented runs each completed 10,240 warmup and 10,240 measured requests at the original load with zero client failures; each build has two V1 and two V2 repetitions.The remaining throughput difference meets the requested acceptance tolerance. Patched modal GEN device times are 4.337 ms at padded batch72 for V1 and 4.581 ms at batch80 for V2; these are different batch sizes, not a same-kernel comparison.
Draft review notes: mean TPOT remains +10.28%; mean-of-round p95/p99 TPOT remain +121.26%/+38.80%. Larger final decode cohorts correlate with the tails, while middle-cohort TPOT remains about +4.9%; CPU/batching attribution is still incomplete. Full-output text differences also remain unresolved without fixed-batch token/logit or reference-quality validation. CTX timeout warnings at the warmup/measurement boundary predate this patch; deferred reaping is supported by source/timing, but per-warning transfer-completion timestamps are unavailable.
PR Checklist
GitHub Bot Help
CI has not been requested for this draft.