Skip to content

Allow independent QKV precision in PyTorch quantization passes - #2700

Open
Ti-Tai Wang (titaiwangms) wants to merge 2 commits into
mainfrom
titaiwang/independent-qkv-quantization
Open

Ti-Tai Wang (titaiwangms) wants to merge 2 commits into
mainfrom
titaiwang/independent-qkv-quantization

Conversation

@titaiwangms

Copy link
Copy Markdown
Contributor

Summary

Add an opt-in independent_qkv: true setting to the native PyTorch RTN, KQuant,
and GPTQ passes. It preserves separate quantization settings for split attention
Q/K/V projections (for example, INT4 Q/K and INT8 V). The default stays false,
so existing packed-QKV workflows retain their current normalization behavior.

The option guards both QKV-normalization points: fresh quantization and
quantization of a partially quantized checkpoint. Existing quantized weights
remain locked, and the option is invocation-local; it is not persisted in the
checkpoint schema. SelectiveMixedPrecision, AutoClip, ONNX exporters, and
quantization kernels are unchanged.

{
  "type": "Rtn",
  "bits": 4,
  "group_size": 128,
  "independent_qkv": true,
  "overrides": {
    "re:.*\\.self_attn\\.v_proj": {"bits": 8}
  }
}

Only use this option with downstream exporters that retain separate Q/K/V
projections; it does not enable mixed-width packed-QKV MatMul.

Validation

  • test_independent_qkv.py exercises real RTN, KQuant, and GPTQ passes, the
    original default, explicit/regex overrides, disk save/reload and packed
    buffers, locked checkpoint merges, and exclusions.
  • A tiny offline Qwen3-MoE RTN checkpoint retains Q/K/V at 4/4/8 after reload
    while fused experts remain quantized.
  • Full test_quant_utils.py plus the new suite: 94 passed.
  • Existing targeted RTN/KQuant/GPTQ checkpoint and MoE tests: 7 passed.
  • Ruff format and targeted Ruff check passed excluding CPY001, which also
    reports on unchanged files in the base checkout; git diff --check passed.

Full 30B ONNX export, ORT CUDA execution, and the separate mixed-width QMoE
full-logit parity gate were not part of this change and are not claimed to pass.

This upstream-head PR supersedes #2697. The follow-up commit removes Pylint W0212 protected-member access in the new tests.

Preserve per-projection RTN, KQuant, and GPTQ overrides behind an explicit opt-in while retaining the existing shared-QKV default for packed-weight exporters. Cover fresh checkpoints, locked checkpoint merges, serialization, and tiny Qwen3-MoE fused experts without changing SMP or downstream export semantics.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: titaiwangms <titaiwang@microsoft.com>
Replace protected pass and module accesses with the public default_config and weight APIs. Preserve the locked-weight and checkpoint assertions while clearing Pylint W0212 in the Python format check.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: titaiwangms <titaiwang@microsoft.com>

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟢 Approval recommended

The implementation is backward-compatible, scoped, documented, and thoroughly tested.

Review effort: Balanced
Findings: None

What changed in this PR

Adds opt-in independent Q/K/V precision to native PyTorch quantization while preserving existing behavior by default.

Changes:

  • Exposes independent_qkv for RTN, KQuant, and GPTQ.
  • Skips QKV normalization when enabled, including checkpoint merges.
  • Adds documentation and comprehensive offline tests.
File Description
olive/​passes/​pytorch/​quant_utils.py Defines and applies the new option.
olive/​passes/​pytorch/​rtn.py Enables the option for RTN.
olive/​passes/​pytorch/​kquant.py Enables the option for KQuant.
olive/​passes/​pytorch/​gptq.py Enables the option for GPTQ.
test/​passes/​pytorch/​test_independent_qkv.py Tests defaults, overrides, reloads, exclusions, and MoE.
docs/​source/​features/​quantization.md Documents usage and exporter limitations.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants