Skip to content

fix(families): repair Community GPU input contracts - #1611

Draft
yifeif-nv wants to merge 5 commits into
NVIDIA:mainfrom
yifeif-nv:fix/community-gpu-ci-reliability-20261008
Draft

yifeif-nv wants to merge 5 commits into
NVIDIA:mainfrom
yifeif-nv:fix/community-gpu-ci-reliability-20261008

Conversation

@yifeif-nv

@yifeif-nv yifeif-nv commented Oct 8, 2026 •

Copy link
Copy Markdown
Collaborator

Background

Community GPU failures exposed inherited family contracts that fail independently of the PR being tested: ConvBERT/DeBERTa bundles omit the serialized tokenizer, DeepSeek-V2 queries Hub metadata during offline testing, and Gemma's size-only budget unit allocates a very large dense array.

Exit Criteria

  • Official slow-format tokenizer assets produce bundles consumable by the native tokenizer.
  • A staged DeepSeek-V2 checkpoint is consumed without an online Hub request.
  • Gemma's original budget assertions execute without a hundred-GiB array allocation.
  • Community explicitly reports the two large Gemma cases outside its scope while preserving their existing premerge declarations and criteria.

Implementation

Each repair remains family-owned and is split into its own commit:

  • Serialize ConvBERT and DeBERTa fast tokenizers directly into the bundle, without modifying read-only checkpoint caches or duplicating sections.
  • Request local-only DeepSeek-V2 checkpoint resolution.
  • Represent Gemma's zero-valued size fixture with a zero-stride array preserving shape, size, dtype, and all existing assertions.
  • Set community_gpu: false for Gemma 3 12B/27B only, retaining premerge: true. The generic manifest guard recognizes and validates this boolean metadata.

Change categories

  • Model or runtime behavior
  • Public API
  • ABI
  • Bundle or artifact format
  • Dependencies
  • Documentation only
  • CI or developer tooling

Validation

Commands and Results

Published head 53a92f5c6ebc15ba9320f7d6a6afb56c0fd47952, based on main 9b083a7fdac56f9d0f14084e61e2aa9dc0dc8816:

  • python -m pytest families/convbert/tests families/deberta/tests families/deepseek_v2/tests families/gemma/tests -q: 94 passed, 15 existing unselected GPU E2Es skipped in a fresh CPU-only repository container.

  • Closed-manifest and premerge guards: 2 passed.

  • Real native WordPiece/BPE consumers read fixed bundles generated from pinned official tokenizer assets: both passed. Weight/engine generation was stubbed in this targeted asset-contract control; it is not model GPU parity.

  • Original DeepSeek/Gemma assertions and numerical criteria were preserved; Ruff and git diff --check passed.

  • Stable CPU run 37753257071: 1,534 Python tests passed, 2 skipped; 246/246 CTests passed. Completed Dev GPU run 37756059191 tested the same head through merge 8900d3b8086eb8eb675b46a7cfebb1df5c99e010 with CI 6e9cc60beafd1b38d0492144c6aa471837a33135: 15/16 selected cases passed on AWS L4. ConvBERT, DeBERTa, all four selected Gemma cases, and the five baseline families passed. The run remains failed because DeepSeek-V2 reached engine construction and TensorRT rejected addMoE on L4; this is no longer an offline checkpoint-resolution failure. The single VM reached functional readiness and deletion was confirmed by both cleanup paths.

  • Unchanged local Blackwell control: E2ERunner(CiContext(repo, env))._run(("deepseek_v2",), ("deepseek-v2-tiny",)): 1 selected E2E passed, 0 failed, 0 selected skipped; 13 family unit tests and 1 CTest passed on GB300. The full checkpoint was pinned to katuni4ka/tiny-random-deepseek-v3@ba144b0d3331a5892aa588d82722d382be2b6e6b; original FP16/TP1/NED <=0.25 criteria were unchanged. This is the existing text-parity contract, not a claim of exact token equality.

Hardware, Environment, and Revisions

Published head 53a92f5c on main 9b083a7f. Local checks used an isolated Linux ARM64 CPU-only container with Python 3.12, TensorRT 11.1.0.106 and Transformers 5.2.0. Native tokenizer controls used pinned official ConvBERT and DeBERTa assets; they did not execute GPU model inference. Completed Dev qualification used Linux x86_64, AWS g6.4xlarge / L4, and TensorRT 11.1.0.106. The unchanged DeepSeek control used Linux ARM64, GB300 (SM10.3), Torch 2.12.0+cu130, Transformers 5.2.0, Hub 1.33.0, NumPy 2.4.6, and tokenizers 0.22.2.

Not Run / Remaining Gaps

The AWS Dev run is not green: DeepSeek-V2 native MoE is outside the documented TensorRT 11.1 MoE capability boundary. Its unchanged tiny case passed locally on GB300, but a qualified Community SM10.x/11.x execution route is still pending. Fifteen unselected local CPU-container GPU cases were skipped; targeted tokenizer asset controls stubbed engine and weight generation. Gemma 12B/27B remain explicitly outside Community coverage and are not qualified by these runs. Other checkpoints, TP modes, Internal/Nightly qualification, and performance are not covered by this evidence.

Contributor Self-Review

  • I have completed a self-review of this change.

Notes For Future Readers

Community routing requires the generic Dev runner support in #1610. It does not remove the two large cases from Internal/Nightly qualification, and deferral is not a passing result for those cases. Review the family serialization/offline fixes before the fixture and routing changes. This PR does not claim that every failure seen in the broad #1053 run was introduced by that PR or is resolved here.

Risk level

  • Low
  • Medium
  • High

The change affects GPU validation or artifact contracts and requires the recorded target-platform qualification; existing numerical criteria and explicit failure gates are retained.

Produce tokenizer.json from the loaded fast tokenizer so slow-format
checkpoints satisfy the native bundle contract. Preserve checkpoint files
and special-token framing, including when tokenizer.json already exists.

Validation: two CPU bundle/tokenization regressions and the native
WordPiece consumer with the pinned official checkpoint assets.

Signed-off-by: yifeif <277870278+yifeif-nv@users.noreply.github.com>
Produce tokenizer.json from the loaded fast tokenizer so slow-format
checkpoints satisfy the native BPE bundle contract. Preserve checkpoint
files and special-token framing without duplicate asset sections.

Validation: two CPU bundle/tokenization regressions and the native BPE
consumer with the pinned official checkpoint assets.

Signed-off-by: yifeif <277870278+yifeif-nv@users.noreply.github.com>
Honor HF_HUB_OFFLINE explicitly when resolving the E2E checkpoint so
a cached pinned revision does not require remote Hub tree metadata.
Keep online lookup and the required config.json assertion unchanged.

Validation: six real Hub-cache contract regressions in the CPU container.
Signed-off-by: yifeif <277870278+yifeif-nv@users.noreply.github.com>
Represent the original zero-valued weight shapes with one backing float.
The size estimator reads metadata only; dense arrays needlessly requested
up to 101 GiB of address space. Preserve every original assertion and
budget threshold.

Validation: all four split-budget tests in the supported CPU container.
Signed-off-by: yifeif <277870278+yifeif-nv@users.noreply.github.com>
Declare the 12B and 27B cases unavailable for the Community GPU resource
envelope. Keep premerge tags, references, and thresholds unchanged for
Internal CI and Nightly. Extend the generic manifest metadata contract
with a validated boolean community_gpu field.

The selected Dev planner consumes this field and reports deferred cases;
this commit does not change the Stable planner or numerical criteria.

Validation: closed-manifest and explicit-premerge architecture contracts.
Signed-off-by: yifeif <277870278+yifeif-nv@users.noreply.github.com>
@coderabbitai

coderabbitai Bot commented Oct 8, 2026

Copy link
Copy Markdown

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Comment @coderabbitai help to get the list of available commands.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant