Skip to content

Allow independent QKV precision in PyTorch quantization passes - #2697

Closed
Ti-Tai Wang (titaiwangms) wants to merge 1 commit into
microsoft:mainfrom
titaiwangms:titaiwang/independent-qkv-quantization
Closed

Ti-Tai Wang (titaiwangms) wants to merge 1 commit into
microsoft:mainfrom
titaiwangms:titaiwang/independent-qkv-quantization

Conversation

@titaiwangms

Copy link
Copy Markdown
Contributor

Summary

Add an opt-in independent_qkv: true setting to the native PyTorch RTN, KQuant,
and GPTQ passes. It preserves separate quantization settings for split attention
Q/K/V projections (for example, INT4 Q/K and INT8 V). The default stays false,
so existing packed-QKV workflows retain their current normalization behavior.

The option guards both QKV-normalization points: fresh quantization and
quantization of a partially quantized checkpoint. Existing quantized weights
remain locked, and the option is invocation-local; it is not persisted in the
checkpoint schema. SelectiveMixedPrecision, AutoClip, ONNX exporters, and
quantization kernels are unchanged.

{
  "type": "Rtn",
  "bits": 4,
  "group_size": 128,
  "independent_qkv": true,
  "overrides": {
    "re:.*\\.self_attn\\.v_proj": {"bits": 8}
  }
}

Only use this option with downstream exporters that retain separate Q/K/V
projections; it does not enable mixed-width packed-QKV MatMul.

Validation

  • test_independent_qkv.py exercises real RTN, KQuant, and GPTQ passes, the
    original default, explicit/regex overrides, disk save/reload and packed
    buffers, locked checkpoint merges, and exclusions.
  • A tiny offline Qwen3-MoE RTN checkpoint retains Q/K/V at 4/4/8 after reload
    while fused experts remain quantized.
  • Full test_quant_utils.py plus the new suite: 94 passed.
  • Existing targeted RTN/KQuant/GPTQ checkpoint and MoE tests: 7 passed.
  • Ruff format and targeted Ruff check passed excluding CPY001, which also
    reports on unchanged files in the base checkout; git diff --check passed.

Full 30B ONNX export, ORT CUDA execution, and the separate mixed-width QMoE
full-logit parity gate were not part of this change and are not claimed to pass.

Preserve per-projection RTN, KQuant, and GPTQ overrides behind an explicit opt-in while retaining the existing shared-QKV default for packed-weight exporters. Cover fresh checkpoints, locked checkpoint merges, serialization, and tiny Qwen3-MoE fused experts without changing SMP or downstream export semantics.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: titaiwangms <titaiwang@microsoft.com>
Copilot AI balanced review requested due to automatic review settings September 29, 2026 01:26

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟢 Approval recommended

The implementation is scoped, backward-compatible, documented, and comprehensively tested.

Review effort: Balanced
Findings: None

What changed in this PR

Adds opt-in independent Q/K/V precision for native PyTorch quantization while preserving existing normalization by default.

Changes:

  • Exposes independent_qkv for RTN, KQuant, and GPTQ.
  • Bypasses QKV normalization when enabled, including checkpoint merges.
  • Adds documentation and comprehensive offline tests.
File Description
olive/​passes/​pytorch/​quant_utils.py Defines and applies the new option.
olive/​passes/​pytorch/​rtn.py Enables the option for RTN.
olive/​passes/​pytorch/​kquant.py Enables the option for KQuant.
olive/​passes/​pytorch/​gptq.py Enables the option for GPTQ.
test/​passes/​pytorch/​test_independent_qkv.py Tests defaults, overrides, reloads, exclusions, and MoE.
docs/​source/​features/​quantization.md Documents usage and exporter limitations.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

for pass_type in (Rtn, KQuant, Gptq):
quantizer = create_pass_from_dict(pass_type, {}, disable_search=True)
assert quantizer.config.independent_qkv is False
assert "independent_qkv" not in AutoClip._default_config(None)
independent_qkv=opt_in,
)
loaded = _assert_qkv(second, path, (4, 4, 8) if opt_in else (8, 8, 8))
assert loaded.get_submodule(V)._parameters["weight"].qweight.equal(original_v.qweight)
modules_to_not_convert=[K],
)
first_loaded = first.load_model()
q_weight = first_loaded.get_submodule(Q)._parameters["weight"].qweight.clone()
)
first_loaded = first.load_model()
q_weight = first_loaded.get_submodule(Q)._parameters["weight"].qweight.clone()
v_weight = first_loaded.get_submodule(V)._parameters["weight"].qweight.clone()
path = tmp_path / "second"
second = _run(pass_type, first, path, bits=4, independent_qkv=True)
loaded = _assert_qkv(second, path, (4, 4, 8))
assert loaded.get_submodule(Q)._parameters["weight"].qweight.equal(q_weight)
second = _run(pass_type, first, path, bits=4, independent_qkv=True)
loaded = _assert_qkv(second, path, (4, 4, 8))
assert loaded.get_submodule(Q)._parameters["weight"].qweight.equal(q_weight)
assert loaded.get_submodule(V)._parameters["weight"].qweight.equal(v_weight)
qcfg = loaded.config.quantization_config
assert_packed_quant_module(loaded.get_submodule(Q), bits=4, group_size=16, symmetric=True)
assert_packed_quant_module(loaded.get_submodule(V), bits=8, group_size=32, symmetric=False)
assert_saved_quant_tensor_matches(path, V, loaded.get_submodule(V)._parameters["weight"])
assert_packed_quant_module(loaded.get_submodule(Q), bits=4, group_size=16, symmetric=True)
assert_packed_quant_module(loaded.get_submodule(V), bits=8, group_size=32, symmetric=False)
assert_saved_quant_tensor_matches(path, V, loaded.get_submodule(V)._parameters["weight"])
assert not hasattr(loaded.get_submodule(K)._parameters["weight"].data, "qweight")
@titaiwangms

Copy link
Copy Markdown
Contributor Author

Superseded by #2700, which uses an upstream microsoft/Olive head branch and includes the Pylint W0212 test fix. Closing this fork-head PR to keep review in one place.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants