Repository navigation
Conversation
…stration ROCm#290 pointed PyTorch>=2.11 at torchao.quantization.pt2e.quantize_pt2e after torch.ao.quantization.quantize_pt2e was removed upstream (confirmed absent on 2.13.0+rocm7.2). That works, but makes torchao a hard new dependency for what is only a registration side effect (populating torch.ops.quantized_decomposed). torch.ao.quantization.fx._decomposed provides the exact same registration with zero extra dependencies -- confirmed live: torch.ops.quantized_decomposed .quantize_per_tensor is absent before this import and present after, on 2.13.0+rocm7.2 (ROCm/MI300X). It's also where quantize_pt2e itself sourced this from on older PyTorch, so it's a strictly more fundamental import, not a workaround. Practical motivation: torchao ships CUDA-specific compiled extensions (_C_mxfp8, _C_cutlass_90a) that fail to load on a ROCm build with no functional impact, which reads as a confusing/alarming warning for anyone who doesn't already depend on torchao for something else -- pure noise from an otherwise-unrelated import. Prefers the dependency-free path; falls back to the existing ROCm#290 behavior (torchao on >=2.11, in-tree quantize_pt2e below that) only if _decomposed is ever unavailable, so this is strictly additive and preserves existing behavior on any PyTorch where the new path doesn't apply. Tested on real MI300X hardware (ROCm 7.2.4, torch 2.13.0+rocm7.2): clean import with no torchao installed, torch.compile(backend="migraphx") compiles and runs correctly (cosine similarity 1.0 vs eager on a transformer-block shaped module).
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
#290 pointed PyTorch >= 2.11 at
torchao.quantization.pt2e.quantize_pt2eaftertorch.ao.quantization.quantize_pt2ewas removed upstream (confirmed absent on 2.13.0+rocm7.2). That works, but makestorchaoa hard new dependency for what is only a registration side effect here (populatingtorch.ops.quantized_decomposed).torch.ao.quantization.fx._decomposedprovides the exact same registration with zero extra dependencies — confirmed live:torch.ops.quantized_decomposed.quantize_per_tensoris absent before this import and present after, on 2.13.0+rocm7.2 (real MI300X). It's also wherequantize_pt2eitself sourced this registration from on older PyTorch, so it's a strictly more fundamental import rather than a workaround.Practical motivation:
torchaoships CUDA-specific compiled extensions (_C_mxfp8,_C_cutlass_90a) that fail to load on a ROCm build with no functional impact — confusing/alarming noise for anyone who doesn't already depend on torchao for something else.This PR prefers the dependency-free path and falls back to the existing #290 behavior (torchao on >=2.11, in-tree
quantize_pt2ebelow that) only if_decomposedis ever unavailable, so it's strictly additive and preserves existing behavior everywhere the new path doesn't apply.Test plan
import torch_migraphxwith notorchaoinstalledtorch.compile(backend="migraphx")compiles and runs correctly on a transformer-block-shaped module: cosine similarity 1.0 vs eager🤖 Generated with Claude Code