Skip to content

[None][fix] Bound V2 disaggregated KV transfer admission - #19855

Draft
yizhang-nv wants to merge 2 commits into
NVIDIA:mainfrom
yizhang-nv:codex/fix-gptoss-disagg-kvcm-v2-perf-20261004
Draft

yizhang-nv wants to merge 2 commits into
NVIDIA:mainfrom
yizhang-nv:codex/fix-gptoss-disagg-kvcm-v2-perf-20261004

Conversation

@yizhang-nv

@yizhang-nv yizhang-nv commented Oct 5, 2026 •

Copy link
Copy Markdown
Member

Description

The asynchronous Python transfer path with KVCM V2 and PP=1 bypasses the executor transfer window. At high disaggregated concurrency, inline receive setup repeatedly waits for bounce-buffer reservations on the GEN execution thread, starving decode launches. GPT-OSS 120B FP4 at ISL8192/OSL1024 and concurrency1024 reproduced a 21.92% output-throughput regression against a measured V1 baseline.

Restore the transfer window and share its budget with the V2 scheduler before KV allocation. Defer new receives without starving active decode, preserve admission order among KV-fitting requests and oversized-request progress, and distinguish transfer backpressure from KV exhaustion. The early path is limited to asynchronous Python V2/PP1; other configurations keep post-scheduling admission. Applying only a late window caused repeated V2 allocation/rollback, which the early budget avoids.

Related: #17495. No public configuration/API, dependency, workload, test selection, or waiver changes.

Test Coverage

Validated on two ARM GB200 nodes with CTX TP1 / GEN TP4 from main fc0876cfd6c5d661f186707857a4e5bfeb12bc6f, using a clean native/bindings/wheel build and verified runtime imports. The Python patch retains identical native artifacts and dependencies. Eight uninstrumented runs each completed 10,240 warmup and 10,240 measured requests at the original load with zero client failures; each build has two V1 and two V2 repetitions.

Build V1 output tok/s V2 output tok/s V2 vs V1 V1 mean TPOT V2 mean TPOT
Original 16,084.10 12,559.23 -21.92% 4.464 ms 18.312 ms
Patched 16,066.51 15,880.01 -1.16% 4.472 ms 4.932 ms

The remaining throughput difference meets the requested acceptance tolerance. Patched modal GEN device times are 4.337 ms at padded batch72 for V1 and 4.581 ms at batch80 for V2; these are different batch sizes, not a same-kernel comparison.

  • The PR adds 11 focused regression cases across three existing test files: early budget enforcement, continued decode, KV-fit ordering, oversized-request progress, budget release, fail-fast handling, and async-path eligibility. Existing admission assertions are updated for the restored window.
  • All retained test functions are unchanged from the completed ARM validation. The product modules retain the exact tested source fingerprints; the follow-up commit only removes broader investigation tests.
  • Separately, the investigation completed 488 distinct targeted tests with zero skips, including native bounce pressure/fallback/reuse, exact KV tensor equality, and deferred-copy ordering. The expanded transfer matrix and test-helper changes are archived locally and are not part of this PR.
  • Those native transfer tests use one process and one physical GPU with logical ranks; they do not independently prove cross-node tensor or full-model semantic parity. The full two-node serving runs are separate evidence.
  • Commit hooks passed. No CI bot run was requested.

Draft review notes: mean TPOT remains +10.28%; mean-of-round p95/p99 TPOT remain +121.26%/+38.80%. Larger final decode cohorts correlate with the tails, while middle-cohort TPOT remains about +4.9%; CPU/batching attribution is still incomplete. Full-output text differences also remain unresolved without fixed-batch token/logit or reference-quality validation. CTX timeout warnings at the warmup/measurement boundary predate this patch; deferred reaping is supported by source/timing, but per-warning transfer-completion timestamps are unavailable.

PR Checklist

  • Description explains the trigger, resulting behavior, and measured limits.
  • Coding guidelines followed; targeted tests cover the added paths.
  • No public API, new dependency, ownership, or significant architecture change.
  • Draft: maintainer/code-owner review and CI are still pending.
  • Reviewed the template items as applicable to this draft.

GitHub Bot Help

CI has not been requested for this draft.

Signed-off-by: Yi Zhang <187001205+yizhang-nv@users.noreply.github.com>
Signed-off-by: Yi Zhang <187001205+yizhang-nv@users.noreply.github.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant