Skip to content

[DSL] Expose F32_FP8 conversion from packed_i32 - #1109

Open
big-yellow-duck wants to merge 11 commits into
ROCm:mainfrom
big-yellow-duck:feat/rocdl-cvt-f32-fp8
Open

big-yellow-duck wants to merge 11 commits into
ROCm:mainfrom
big-yellow-duck:feat/rocdl-cvt-f32-fp8

Conversation

@big-yellow-duck

Copy link
Copy Markdown
Contributor

Summary

  • add a typed fx.rocdl.cvt_f32_fp8 wrapper for the scalar ROCDL conversion
  • export it through the public ROCDL expression API
  • add a backend-agnostic test for the emitted operation and result type

Motivation

RDNA4 SplitKV needs scalar FP8 conversion for irregular key/value loads where a packed conversion is not applicable. The underlying MLIR ROCDL operation already exists, but FlyDSL does not currently expose it.

API stability

This is an additive stable API change. The stable API catalog gains flydsl.expr.rocdl.cvt_f32_fp8; no existing entries are removed or changed.

Validation

  • bash scripts/check_python_style.sh
  • python3 -m pytest tests/unit/test_rocdl_conversions.py -q
  • stable API catalog comparison against main

big-yellow-duck and others added 4 commits September 9, 2026 06:32
Add a typed ROCDL wrapper for selecting and decoding one E4M3 FP8 byte from a
packed i32 source. Export the wrapper through the stable backend API and verify
its result type and emitted operation.

Co-authored-by: vllmellm <190700713+vllmellm@users.noreply.github.com>
Co-authored-by: Tan Pin Siang <1716735+tanpinsiang@users.noreply.github.com>
Signed-off-by: big_yellow_duck <83417790+big-yellow-duck@users.noreply.github.com>
@coderfeli

Copy link
Copy Markdown
Collaborator

why not Float32() directly?

@big-yellow-duck

Copy link
Copy Markdown
Contributor Author

Float32() works when its operand is already a typed FP8 value, but this API is needed for packed raw data.

The consumer is the SplitKV kernel in vLLM PR #55996. It performs a 64-bit copy into an eight-element Uint8 register fragment, then loads that fragment and dequantizes its eight FP8 payloads.

The helper bitcasts those eight bytes into two packed i32 words and decodes bytes 0–3 from each word. This matches the ROCDL/hardware interface: an i32 source plus a byte selector.

Calling Float32(packed_i32) would instead perform a numerical integer-to-float conversion of the entire word; it would not select and interpret one byte as FP8. We could conceptually bitcast the bytes to Float8E4M3FN first and then convert to Float32, but the current FlyDSL ROCm pipeline does not lower that FP8 arith.extf path to this ROCDL instruction.

IF there was an api to do Float32(packged_i32) and compile down to cvt_f32_fp8 instuction that would be best.

@big-yellow-duck big-yellow-duck changed the title [DSL] Expose scalar FP8 conversion [DSL] Expose F32_FP8 conversion from packed_i32 Sep 12, 2026
@big-yellow-duck

Copy link
Copy Markdown
Contributor Author

Hi @coderfeli is it all ok from your side? hope we can get this merged into the next release

prcoe1 added a commit to prcoe1/r9700-serving that referenced this pull request Sep 27, 2026
…experiment

- Remove 55996-rdna4-splitkv.patch + v030-port: ROCM_ATTN/Triton-SplitKV
  A/B vs tuned UA lost everywhere measured (pp2048 -9~10%, tg32 -6%,
  tg128 -9% at d0); FlyDSL fast path needs unmerged ROCm/FlyDSL#1109.
  Keep #56005 (live GEMM) and #57767 (dormant, unrelated).
- qwen3.8-27b.env: c4 with spec-decode disabled (repeat of the 09-23
  throughput-first shape at conc 4); #35288/#55533 triggers are MTP-only.
- compose.yaml backend back to hardcoded UA; README SplitKV entry out.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants