Allow independent QKV precision in PyTorch quantization passes - #2697
Closed
Ti-Tai Wang (titaiwangms) wants to merge 1 commit into
Closed
Ti-Tai Wang (titaiwangms) wants to merge 1 commit into
Ti-Tai Wang (titaiwangms) wants to merge 1 commit into
Conversation
Preserve per-projection RTN, KQuant, and GPTQ overrides behind an explicit opt-in while retaining the existing shared-QKV default for packed-weight exporters. Cover fresh checkpoints, locked checkpoint merges, serialization, and tiny Qwen3-MoE fused experts without changing SMP or downstream export semantics. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: titaiwangms <titaiwang@microsoft.com>
Copilot started reviewing on behalf of
Ti-Tai Wang (titaiwangms)
September 29, 2026 01:27
View session
Contributor
There was a problem hiding this comment.
Copilot review overview
🟢 Approval recommended
The implementation is scoped, backward-compatible, documented, and comprehensively tested.
Review effort: Balanced
Findings: None
What changed in this PR
Adds opt-in independent Q/K/V precision for native PyTorch quantization while preserving existing normalization by default.
Changes:
- Exposes
independent_qkvfor RTN, KQuant, and GPTQ. - Bypasses QKV normalization when enabled, including checkpoint merges.
- Adds documentation and comprehensive offline tests.
| File | Description |
|---|---|
olive/passes/pytorch/quant_utils.py |
Defines and applies the new option. |
olive/passes/pytorch/rtn.py |
Enables the option for RTN. |
olive/passes/pytorch/kquant.py |
Enables the option for KQuant. |
olive/passes/pytorch/gptq.py |
Enables the option for GPTQ. |
test/passes/pytorch/test_independent_qkv.py |
Tests defaults, overrides, reloads, exclusions, and MoE. |
docs/source/features/quantization.md |
Documents usage and exporter limitations. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
| for pass_type in (Rtn, KQuant, Gptq): | ||
| quantizer = create_pass_from_dict(pass_type, {}, disable_search=True) | ||
| assert quantizer.config.independent_qkv is False | ||
| assert "independent_qkv" not in AutoClip._default_config(None) |
| independent_qkv=opt_in, | ||
| ) | ||
| loaded = _assert_qkv(second, path, (4, 4, 8) if opt_in else (8, 8, 8)) | ||
| assert loaded.get_submodule(V)._parameters["weight"].qweight.equal(original_v.qweight) |
| modules_to_not_convert=[K], | ||
| ) | ||
| first_loaded = first.load_model() | ||
| q_weight = first_loaded.get_submodule(Q)._parameters["weight"].qweight.clone() |
| ) | ||
| first_loaded = first.load_model() | ||
| q_weight = first_loaded.get_submodule(Q)._parameters["weight"].qweight.clone() | ||
| v_weight = first_loaded.get_submodule(V)._parameters["weight"].qweight.clone() |
| path = tmp_path / "second" | ||
| second = _run(pass_type, first, path, bits=4, independent_qkv=True) | ||
| loaded = _assert_qkv(second, path, (4, 4, 8)) | ||
| assert loaded.get_submodule(Q)._parameters["weight"].qweight.equal(q_weight) |
| second = _run(pass_type, first, path, bits=4, independent_qkv=True) | ||
| loaded = _assert_qkv(second, path, (4, 4, 8)) | ||
| assert loaded.get_submodule(Q)._parameters["weight"].qweight.equal(q_weight) | ||
| assert loaded.get_submodule(V)._parameters["weight"].qweight.equal(v_weight) |
| qcfg = loaded.config.quantization_config | ||
| assert_packed_quant_module(loaded.get_submodule(Q), bits=4, group_size=16, symmetric=True) | ||
| assert_packed_quant_module(loaded.get_submodule(V), bits=8, group_size=32, symmetric=False) | ||
| assert_saved_quant_tensor_matches(path, V, loaded.get_submodule(V)._parameters["weight"]) |
| assert_packed_quant_module(loaded.get_submodule(Q), bits=4, group_size=16, symmetric=True) | ||
| assert_packed_quant_module(loaded.get_submodule(V), bits=8, group_size=32, symmetric=False) | ||
| assert_saved_quant_tensor_matches(path, V, loaded.get_submodule(V)._parameters["weight"]) | ||
| assert not hasattr(loaded.get_submodule(K)._parameters["weight"].data, "qweight") |
Contributor
Author
|
Superseded by #2700, which uses an upstream microsoft/Olive head branch and includes the Pylint W0212 test fix. Closing this fork-head PR to keep review in one place. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Add an opt-in
independent_qkv: truesetting to the native PyTorch RTN, KQuant,and GPTQ passes. It preserves separate quantization settings for split attention
Q/K/V projections (for example, INT4 Q/K and INT8 V). The default stays
false,so existing packed-QKV workflows retain their current normalization behavior.
The option guards both QKV-normalization points: fresh quantization and
quantization of a partially quantized checkpoint. Existing quantized weights
remain locked, and the option is invocation-local; it is not persisted in the
checkpoint schema. SelectiveMixedPrecision, AutoClip, ONNX exporters, and
quantization kernels are unchanged.
{ "type": "Rtn", "bits": 4, "group_size": 128, "independent_qkv": true, "overrides": { "re:.*\\.self_attn\\.v_proj": {"bits": 8} } }Only use this option with downstream exporters that retain separate Q/K/V
projections; it does not enable mixed-width packed-QKV MatMul.
Validation
test_independent_qkv.pyexercises real RTN, KQuant, and GPTQ passes, theoriginal default, explicit/regex overrides, disk save/reload and packed
buffers, locked checkpoint merges, and exclusions.
while fused experts remain quantized.
test_quant_utils.pyplus the new suite: 94 passed.CPY001, which alsoreports on unchanged files in the base checkout;
git diff --checkpassed.Full 30B ONNX export, ORT CUDA execution, and the separate mixed-width QMoE
full-logit parity gate were not part of this change and are not claimed to pass.