Allow independent QKV precision in PyTorch quantization passes - #2700
Open
Ti-Tai Wang (titaiwangms) wants to merge 2 commits into
Open
Ti-Tai Wang (titaiwangms) wants to merge 2 commits into
Ti-Tai Wang (titaiwangms) wants to merge 2 commits into
Conversation
Preserve per-projection RTN, KQuant, and GPTQ overrides behind an explicit opt-in while retaining the existing shared-QKV default for packed-weight exporters. Cover fresh checkpoints, locked checkpoint merges, serialization, and tiny Qwen3-MoE fused experts without changing SMP or downstream export semantics. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: titaiwangms <titaiwang@microsoft.com>
Replace protected pass and module accesses with the public default_config and weight APIs. Preserve the locked-weight and checkpoint assertions while clearing Pylint W0212 in the Python format check. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: titaiwangms <titaiwang@microsoft.com>
Copilot started reviewing on behalf of
Ti-Tai Wang (titaiwangms)
September 29, 2026 23:51
View session
Contributor
There was a problem hiding this comment.
Copilot review overview
🟢 Approval recommended
The implementation is backward-compatible, scoped, documented, and thoroughly tested.
Review effort: Balanced
Findings: None
What changed in this PR
Adds opt-in independent Q/K/V precision to native PyTorch quantization while preserving existing behavior by default.
Changes:
- Exposes
independent_qkvfor RTN, KQuant, and GPTQ. - Skips QKV normalization when enabled, including checkpoint merges.
- Adds documentation and comprehensive offline tests.
| File | Description |
|---|---|
olive/passes/pytorch/quant_utils.py |
Defines and applies the new option. |
olive/passes/pytorch/rtn.py |
Enables the option for RTN. |
olive/passes/pytorch/kquant.py |
Enables the option for KQuant. |
olive/passes/pytorch/gptq.py |
Enables the option for GPTQ. |
test/passes/pytorch/test_independent_qkv.py |
Tests defaults, overrides, reloads, exclusions, and MoE. |
docs/source/features/quantization.md |
Documents usage and exporter limitations. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Add an opt-in
independent_qkv: truesetting to the native PyTorch RTN, KQuant,and GPTQ passes. It preserves separate quantization settings for split attention
Q/K/V projections (for example, INT4 Q/K and INT8 V). The default stays
false,so existing packed-QKV workflows retain their current normalization behavior.
The option guards both QKV-normalization points: fresh quantization and
quantization of a partially quantized checkpoint. Existing quantized weights
remain locked, and the option is invocation-local; it is not persisted in the
checkpoint schema. SelectiveMixedPrecision, AutoClip, ONNX exporters, and
quantization kernels are unchanged.
{ "type": "Rtn", "bits": 4, "group_size": 128, "independent_qkv": true, "overrides": { "re:.*\\.self_attn\\.v_proj": {"bits": 8} } }Only use this option with downstream exporters that retain separate Q/K/V
projections; it does not enable mixed-width packed-QKV MatMul.
Validation
test_independent_qkv.pyexercises real RTN, KQuant, and GPTQ passes, theoriginal default, explicit/regex overrides, disk save/reload and packed
buffers, locked checkpoint merges, and exclusions.
while fused experts remain quantized.
test_quant_utils.pyplus the new suite: 94 passed.CPY001, which alsoreports on unchanged files in the base checkout;
git diff --checkpassed.Full 30B ONNX export, ORT CUDA execution, and the separate mixed-width QMoE
full-logit parity gate were not part of this change and are not claimed to pass.
This upstream-head PR supersedes #2697. The follow-up commit removes Pylint W0212 protected-member access in the new tests.