Skip to content

Reject global->LDS direct loads on gfx11 - #1122

Merged
sjfeng1999 merged 1 commit into
mainfrom
matthias.reject-lds-dma-on-gfx11
Sep 22, 2026
Merged

sjfeng1999 merged 1 commit into
mainfrom
matthias.reject-lds-dma-on-gfx11

Conversation

@mgehre-amd

Copy link
Copy Markdown
Contributor

Problem

rocdl.raw.ptr.buffer.load.lds lowers to buffer_load_* ... lds, which has no hardware backing on gfx11 (RDNA3 / RDNA3.5). Emitting it there passes MLIR verification and then aborts the whole process in LLVM instruction selection:

ExpandIntegerOperand Op #2: t171: ch = llvm.amdgcn.raw.ptr.buffer.load.lds<
  (dereferenceable load (s128) from %ir.92, align 1, addrspace 8),
  (dereferenceable store (s4096) into %ir.185, align 1, addrspace 3)> ...

LLVM ERROR: Do not know how to expand this operator's operand!

LLVM ERROR is report_fatal_error → abort() inside the in-process JIT, so the Python interpreter dies with SIGABRT.

This is reachable today: aiter/ops/flydsl/kernels/splitk_hgemm.py selects the async global→LDS path for every arch except gfx942 (ASYNC_COPY = GPU_ARCH != "gfx942"), so running its test on a Strix Halo (gfx1151) kills the interpreter.

Fix

Gate the two entry points that emit the op, so the failure surfaces as a ValueError while a Python traceback still points at the kernel line:

ValueError: raw_ptr_buffer_load_lds is not supported on target arch 'gfx1151': gfx11
(RDNA3 / RDNA3.5) has no global -> LDS direct-load hardware. Use a buffer load into
registers followed by ds_write instead.

Wording follows the existing s_waitcnt "is not supported on target arch …" idiom in the same package, and the precedent set by BufferCopyLDS64b(), which already raises for the same class of bug (verifies, then silently fails instruction selection).

`buffer_load_* ... lds` has no hardware backing on gfx11 (RDNA3 / RDNA3.5);
the path exists on gfx9 and gfx10. Emitting `rocdl.raw.ptr.buffer.load.lds`
there passes MLIR verification and then aborts the whole process in LLVM
instruction selection with "LLVM ERROR: Do not know how to expand this
operator's operand!", killing the Python interpreter with no traceback.

Gate the two entry points that emit the op so the failure surfaces as a
ValueError pointing at the kernel line that asked for it.

Changes:
- The gate reads `env.compile.arch` (`ARCH`) before falling back to
  `get_rocm_arch()`, matching `RocmBackend.detect_target()`. Keying off the
  device instead would accept or reject the wrong target when cross-compiling,
  and it also lets the tests run without a GPU. Neighbouring dispatches in
  universal.py still use bare `get_rocm_arch()`; migrating them is out of scope.
- The helper lives in rocdl/utils.py, the module the sibling arch files already
  pull shared helpers from, so both call sites import it by name.
- `BufferCopyLDS32b` / `128b` now route through `BufferCopyLDS()` instead of
  calling `CopyOpCDNA3BufferCopyLDSType.get()` directly, so they cannot bypass
  the check. That also covers the `fly.copy_atom_call` path into
  `CopyOpCDNA3BufferCopyLDSType::emitAtomCall`, which is unguarded in C++.
- Scope is gfx11 only. gfx1201 and gfx1250 measurably abort the same way and
  are knowingly left alone here; gfx12 wants its own answer (TDM on gfx1250).
- Not covered: `cdna4.BufferLoadAsyncLDS` / `GlobalLoadAsyncLDS`. Same hazard
  class, explicitly gfx950-namespaced, untested here.

Arch classification verified with clang codegen of the bare intrinsic (ROCm
7.15): gfx942 / gfx1010 / gfx1030 select `buffer_load_dword ... lds` with the
`s_mov_b32 m0` setup, while gfx1100 hits report_fatal_error inside
DAGTypeLegalizer::ExpandIntegerOperand. Compiling a minimal FlyDSL kernel per
arch matches. gfx9 / gfx10 are compile-verified only; no gfx10 hardware was
available to run on.
Copilot AI lite review requested due to automatic review settings September 11, 2026 10:44
@mgehre-amd
mgehre-amd requested review from sjfeng1999 and removed request for Copilot September 11, 2026 10:47
@mgehre-amd

Copy link
Copy Markdown
Contributor Author

@sjfeng1999, friendly ping for review (or would someone else be the right reviewer?)

@sjfeng1999

Copy link
Copy Markdown
Collaborator

Thanks for the check. Longer term, adding checks for individual ops doesn’t scale well, since many ops or variants have arch-specific requirements. I’m considering whether we should introduce a systematic framework to annotate each op with its architecture requirements.

@sjfeng1999
sjfeng1999 merged commit 0627ccd into main Sep 22, 2026
17 checks passed
@sjfeng1999
sjfeng1999 deleted the matthias.reject-lds-dma-on-gfx11 branch September 22, 2026 13:34
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants