Skip to content

[Kernel] Refactor preshuffle GEMM indexing with layout algebra - #1131

Merged
coderfeli merged 1 commit into
mainfrom
codex/port-preshuffle-layout-algebra
Sep 23, 2026
Merged

coderfeli merged 1 commit into
mainfrom
codex/port-preshuffle-layout-algebra

Conversation

@coderfeli

Copy link
Copy Markdown
Collaborator

Summary

  • add a shared layout-algebra description for cooperative preshuffle A DMA
  • replace manual DMA, LDS-read, scale, epilogue, C-store, and split-K reduction index arithmetic in the MFMA preshuffle GEMMs
  • represent the DMA mapping with basis strides and CoordSwizzle, avoiding repeated div/mul address calculations in unrolled loads
  • leave rdna_fp8_preshuffle_gemm.py unchanged

Performance

Measured on gfx950 with HIP events, 50 warmups and 500 iterations. Lower latency is better.

path this PR baseline change
FP8 preshuffle 241.2 us 240.9 us +0.1%
MXFP4 55.8 us 58.6 us -4.8%
MXFP6 x FP4 155.2 us 155.9 us -0.4%
MXFP8 102.0 us 108.1 us -5.6%
MXFP8 blockscale 78.9 us 82.9 us -4.8%

Testing

  • ruff check kernels/gemm/preshuffle_layout.py kernels/gemm/preshuffle_gemm.py kernels/gemm/mxfp4_preshuffle.py
  • pytest -q tests/unit/test_layout_algebra.py (31 passed, 1 skipped)
  • targeted gfx950 GPU coverage for FP8 async DMA, MXFP4, MXFP6 x FP4, MXFP8, blockscale, ragged M, fp16/bf16 DMA tile widths, fused epilogues, and CUDA graph capture
  • post-rebase gfx950 smoke: 4 passed

@coderfeli
coderfeli force-pushed the codex/port-preshuffle-layout-algebra branch from 7c8ec0b to f4216b6 Compare September 14, 2026 10:57
@coderfeli
coderfeli merged commit e70bf79 into main Sep 23, 2026
25 of 27 checks passed
@coderfeli
coderfeli deleted the codex/port-preshuffle-layout-algebra branch September 23, 2026 07:58
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant