Repository navigation
Conversation
Produce tokenizer.json from the loaded fast tokenizer so slow-format checkpoints satisfy the native bundle contract. Preserve checkpoint files and special-token framing, including when tokenizer.json already exists. Validation: two CPU bundle/tokenization regressions and the native WordPiece consumer with the pinned official checkpoint assets. Signed-off-by: yifeif <277870278+yifeif-nv@users.noreply.github.com>
Produce tokenizer.json from the loaded fast tokenizer so slow-format checkpoints satisfy the native BPE bundle contract. Preserve checkpoint files and special-token framing without duplicate asset sections. Validation: two CPU bundle/tokenization regressions and the native BPE consumer with the pinned official checkpoint assets. Signed-off-by: yifeif <277870278+yifeif-nv@users.noreply.github.com>
Honor HF_HUB_OFFLINE explicitly when resolving the E2E checkpoint so a cached pinned revision does not require remote Hub tree metadata. Keep online lookup and the required config.json assertion unchanged. Validation: six real Hub-cache contract regressions in the CPU container. Signed-off-by: yifeif <277870278+yifeif-nv@users.noreply.github.com>
Represent the original zero-valued weight shapes with one backing float. The size estimator reads metadata only; dense arrays needlessly requested up to 101 GiB of address space. Preserve every original assertion and budget threshold. Validation: all four split-budget tests in the supported CPU container. Signed-off-by: yifeif <277870278+yifeif-nv@users.noreply.github.com>
Declare the 12B and 27B cases unavailable for the Community GPU resource envelope. Keep premerge tags, references, and thresholds unchanged for Internal CI and Nightly. Extend the generic manifest metadata contract with a validated boolean community_gpu field. The selected Dev planner consumes this field and reports deferred cases; this commit does not change the Stable planner or numerical criteria. Validation: closed-manifest and explicit-premerge architecture contracts. Signed-off-by: yifeif <277870278+yifeif-nv@users.noreply.github.com>
|
Important Draft PR not reviewedDraft PRs are not automatically reviewed by default.
To automatically review draft PRs, update your CodeRabbit configuration: reviews:
auto_review:
drafts: true
Comment |
4 of 11 tasks
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Background
Community GPU failures exposed inherited family contracts that fail independently of the PR being tested: ConvBERT/DeBERTa bundles omit the serialized tokenizer, DeepSeek-V2 queries Hub metadata during offline testing, and Gemma's size-only budget unit allocates a very large dense array.
Exit Criteria
Implementation
Each repair remains family-owned and is split into its own commit:
community_gpu: falsefor Gemma 3 12B/27B only, retainingpremerge: true. The generic manifest guard recognizes and validates this boolean metadata.Change categories
Validation
Commands and Results
Published head
53a92f5c6ebc15ba9320f7d6a6afb56c0fd47952, based on main9b083a7fdac56f9d0f14084e61e2aa9dc0dc8816:python -m pytest families/convbert/tests families/deberta/tests families/deepseek_v2/tests families/gemma/tests -q: 94 passed, 15 existing unselected GPU E2Es skipped in a fresh CPU-only repository container.Closed-manifest and premerge guards: 2 passed.
Real native WordPiece/BPE consumers read fixed bundles generated from pinned official tokenizer assets: both passed. Weight/engine generation was stubbed in this targeted asset-contract control; it is not model GPU parity.
Original DeepSeek/Gemma assertions and numerical criteria were preserved; Ruff and
git diff --checkpassed.Stable CPU run 37753257071: 1,534 Python tests passed, 2 skipped; 246/246 CTests passed. Completed Dev GPU run 37756059191 tested the same head through merge
8900d3b8086eb8eb675b46a7cfebb1df5c99e010with CI6e9cc60beafd1b38d0492144c6aa471837a33135: 15/16 selected cases passed on AWS L4. ConvBERT, DeBERTa, all four selected Gemma cases, and the five baseline families passed. The run remains failed because DeepSeek-V2 reached engine construction and TensorRT rejectedaddMoEon L4; this is no longer an offline checkpoint-resolution failure. The single VM reached functional readiness and deletion was confirmed by both cleanup paths.Unchanged local Blackwell control:
E2ERunner(CiContext(repo, env))._run(("deepseek_v2",), ("deepseek-v2-tiny",)): 1 selected E2E passed, 0 failed, 0 selected skipped; 13 family unit tests and 1 CTest passed on GB300. The full checkpoint was pinned tokatuni4ka/tiny-random-deepseek-v3@ba144b0d3331a5892aa588d82722d382be2b6e6b; original FP16/TP1/NED <=0.25 criteria were unchanged. This is the existing text-parity contract, not a claim of exact token equality.Hardware, Environment, and Revisions
Published head
53a92f5con main9b083a7f. Local checks used an isolated Linux ARM64 CPU-only container with Python 3.12, TensorRT 11.1.0.106 and Transformers 5.2.0. Native tokenizer controls used pinned official ConvBERT and DeBERTa assets; they did not execute GPU model inference. Completed Dev qualification used Linux x86_64, AWSg6.4xlarge/ L4, and TensorRT 11.1.0.106. The unchanged DeepSeek control used Linux ARM64, GB300 (SM10.3), Torch 2.12.0+cu130, Transformers 5.2.0, Hub 1.33.0, NumPy 2.4.6, and tokenizers 0.22.2.Not Run / Remaining Gaps
The AWS Dev run is not green: DeepSeek-V2 native MoE is outside the documented TensorRT 11.1 MoE capability boundary. Its unchanged tiny case passed locally on GB300, but a qualified Community SM10.x/11.x execution route is still pending. Fifteen unselected local CPU-container GPU cases were skipped; targeted tokenizer asset controls stubbed engine and weight generation. Gemma 12B/27B remain explicitly outside Community coverage and are not qualified by these runs. Other checkpoints, TP modes, Internal/Nightly qualification, and performance are not covered by this evidence.
Contributor Self-Review
Notes For Future Readers
Community routing requires the generic Dev runner support in #1610. It does not remove the two large cases from Internal/Nightly qualification, and deferral is not a passing result for those cases. Review the family serialization/offline fixes before the fixture and routing changes. This PR does not claim that every failure seen in the broad #1053 run was introduced by that PR or is resolved here.
Risk level
The change affects GPU validation or artifact contracts and requires the recorded target-platform qualification; existing numerical criteria and explicit failure gates are retained.