Skip to content

test(qwen): reject prompt token count mismatches - #1625

Closed
davemichael wants to merge 1 commit into
NVIDIA:mainfrom
davemichael:test/qwen-prompt-token-parity
Closed

davemichael wants to merge 1 commit into
NVIDIA:mainfrom
davemichael:test/qwen-prompt-token-parity

Conversation

@davemichael

Copy link
Copy Markdown

Background

Qwen3-4B-Instruct-2507 could pass the generation E2E with different native and reference prompts: native execution received 24 tokens while the reference received 20, but the generated tokens matched. The renderer fix is already covered by #1617; this PR adds the missing E2E validation. Related to #1615.

Exit Criteria

Every Qwen generation E2E checks the native prompt token count against the reference-rendered prompt before comparing output. Missing receipts and a mismatch from any reported tensor-parallel rank or repeated run fail. Existing numerical and output acceptance thresholds remain unchanged.

Implementation

Use the reference chat-template rendering to count prompt tokens. Validate native prefill receipts for regular, sampled and repeated generation. Accept MPI-tagged receipts and check each reported rank separately. Changes are confined to three Qwen test files; no runtime, API, ABI, bundle or dependency changes.

Change categories

  • Model or runtime behavior
  • Public API
  • ABI
  • Bundle or artifact format
  • Dependencies
  • Documentation only
  • CI or developer tooling

Validation

Commands and Results

  • python -m pytest families/qwen/tests/test_runtime_receipt.py families/qwen/tests/test_builder_policy.py -q: 13 passed, including rejection of the 24-versus-20-token regression, missing receipts and per-rank mismatches.
  • python -m ruff check families/qwen/tests/runtime_receipt.py families/qwen/tests/test_runtime_receipt.py families/qwen/tests/test_e2e.py: passed.
  • PYTHONPATH=core/builder:apps/benchmark:. python -m tools.model_ci validate: passed.
  • PYTHONPATH=core/builder:apps/benchmark:. python tools/test_impact.py --validate: passed.
  • git diff --check github/main...HEAD: passed. Applicable pre-commit checks passed.

Hardware, Environment, and Revisions

Tested source: 943a07a01af2f5b14af8633c10f3a6dd8b8c029a, based on e6c674e of main. CPU-only checks in a Linux x86_64 development container with Python 3.12 and the pinned requirements/community-ci.txt tools, including pytest 8.4.2 and Ruff 0.16.4. Unit tests use synthetic runtime receipts; no model checkpoint or GPU is needed for these checks.

Not Run / Remaining Gaps

GPU generation, multi-GPU execution, full model parity, performance and protected premerge were not run for this PR head. The Qwen3-4B E2E is expected to expose the known mismatch until #1617 or an equivalent renderer fix lands. Keep this PR in draft pending that dependency and CI.

Contributor Self-Review

  • I have completed a self-review of this change.

Manually reviewed the complete three-file diff at 943a07a01af2f5b14af8633c10f3a6dd8b8c029a, including ordinary/repeated generation, the embedding bypass, reference chat-template counting and tagged rank parsing. No blocking findings within this scope; model execution remains unverified on this head.

Notes For Future Readers

Merge the renderer fix in #1617 first. This PR deliberately avoids duplicating its runtime implementation and native template tests. Equal token counts catch the observed defect but do not establish arbitrary token-by-token prompt identity; missing ranks are not detected by the count check. No new third-party implementation is incorporated.

Risk level

  • Low
  • Medium
  • High

Risk rationale: the stricter E2E check may expose existing prompt mismatches or missing receipts in model configurations that previously passed output-only comparison. Runtime behavior and acceptance tolerances are unchanged.

Compare native prompt receipts with the reference chat template before generation output checks. Validate each reported tensor-parallel rank and repeated run. Complements the renderer fix in NVIDIA#1617.

Co-Authored-By: Codex
Signed-off-by: Dave Michael <dmichael@nvidia.com>
@coderabbitai

coderabbitai Bot commented Oct 8, 2026

Copy link
Copy Markdown

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Comment @coderabbitai help to get the list of available commands.

@davemichael davemichael closed this Oct 8, 2026
@davemichael

Copy link
Copy Markdown
Author

My agent discovered the same bug as #1617
I think the fix and testing there are probably sufficient, so closed this one.

@davemichael
davemichael deleted the test/qwen-prompt-token-parity branch October 8, 2026 21:40
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant