Skip to content

Avoid mandatory torchao dependency for quantized_decomposed registration - #292

Open
npow wants to merge 1 commit into
ROCm:masterfrom
npow:fix/avoid-torchao-dependency-for-quantized-decomposed-registration
Open

npow wants to merge 1 commit into
ROCm:masterfrom
npow:fix/avoid-torchao-dependency-for-quantized-decomposed-registration

Conversation

@npow

@npow npow commented Aug 27, 2026

Copy link
Copy Markdown

Summary

#290 pointed PyTorch >= 2.11 at torchao.quantization.pt2e.quantize_pt2e after torch.ao.quantization.quantize_pt2e was removed upstream (confirmed absent on 2.13.0+rocm7.2). That works, but makes torchao a hard new dependency for what is only a registration side effect here (populating torch.ops.quantized_decomposed).

torch.ao.quantization.fx._decomposed provides the exact same registration with zero extra dependencies — confirmed live: torch.ops.quantized_decomposed.quantize_per_tensor is absent before this import and present after, on 2.13.0+rocm7.2 (real MI300X). It's also where quantize_pt2e itself sourced this registration from on older PyTorch, so it's a strictly more fundamental import rather than a workaround.

Practical motivation: torchao ships CUDA-specific compiled extensions (_C_mxfp8, _C_cutlass_90a) that fail to load on a ROCm build with no functional impact — confusing/alarming noise for anyone who doesn't already depend on torchao for something else.

This PR prefers the dependency-free path and falls back to the existing #290 behavior (torchao on >=2.11, in-tree quantize_pt2e below that) only if _decomposed is ever unavailable, so it's strictly additive and preserves existing behavior everywhere the new path doesn't apply.

Test plan

  • Real MI300X hardware (ROCm 7.2.4, torch 2.13.0+rocm7.2): clean import torch_migraphx with no torchao installed
  • torch.compile(backend="migraphx") compiles and runs correctly on a transformer-block-shaped module: cosine similarity 1.0 vs eager

🤖 Generated with Claude Code

…stration

ROCm#290 pointed PyTorch>=2.11 at torchao.quantization.pt2e.quantize_pt2e after
torch.ao.quantization.quantize_pt2e was removed upstream (confirmed absent on
2.13.0+rocm7.2). That works, but makes torchao a hard new dependency for what
is only a registration side effect (populating torch.ops.quantized_decomposed).

torch.ao.quantization.fx._decomposed provides the exact same registration
with zero extra dependencies -- confirmed live: torch.ops.quantized_decomposed
.quantize_per_tensor is absent before this import and present after, on
2.13.0+rocm7.2 (ROCm/MI300X). It's also where quantize_pt2e itself sourced
this from on older PyTorch, so it's a strictly more fundamental import, not a
workaround.

Practical motivation: torchao ships CUDA-specific compiled extensions
(_C_mxfp8, _C_cutlass_90a) that fail to load on a ROCm build with no
functional impact, which reads as a confusing/alarming warning for anyone who
doesn't already depend on torchao for something else -- pure noise from an
otherwise-unrelated import.

Prefers the dependency-free path; falls back to the existing ROCm#290 behavior
(torchao on >=2.11, in-tree quantize_pt2e below that) only if _decomposed is
ever unavailable, so this is strictly additive and preserves existing
behavior on any PyTorch where the new path doesn't apply.

Tested on real MI300X hardware (ROCm 7.2.4, torch 2.13.0+rocm7.2): clean
import with no torchao installed, torch.compile(backend="migraphx") compiles
and runs correctly (cosine similarity 1.0 vs eager on a transformer-block
shaped module).
@npow
npow requested a review from shivadbhavsar as a code owner August 27, 2026 05:22

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant