Repository navigation
[DSL] Expose F32_FP8 conversion from packed_i32 - #1109
big-yellow-duck wants to merge 11 commits into
Conversation
Add a typed ROCDL wrapper for selecting and decoding one E4M3 FP8 byte from a packed i32 source. Export the wrapper through the stable backend API and verify its result type and emitted operation. Co-authored-by: vllmellm <190700713+vllmellm@users.noreply.github.com> Co-authored-by: Tan Pin Siang <1716735+tanpinsiang@users.noreply.github.com> Signed-off-by: big_yellow_duck <83417790+big-yellow-duck@users.noreply.github.com>
|
why not Float32() directly? |
|
The consumer is the SplitKV kernel in vLLM PR #55996. It performs a 64-bit copy into an eight-element The helper bitcasts those eight bytes into two packed Calling IF there was an api to do |
|
Hi @coderfeli is it all ok from your side? hope we can get this merged into the next release |
…experiment - Remove 55996-rdna4-splitkv.patch + v030-port: ROCM_ATTN/Triton-SplitKV A/B vs tuned UA lost everywhere measured (pp2048 -9~10%, tg32 -6%, tg128 -9% at d0); FlyDSL fast path needs unmerged ROCm/FlyDSL#1109. Keep #56005 (live GEMM) and #57767 (dormant, unrelated). - qwen3.8-27b.env: c4 with spec-decode disabled (repeat of the 09-23 throughput-first shape at conc 4); #35288/#55533 triggers are MTP-only. - compose.yaml backend back to hardcoded UA; README SplitKV entry out.
Summary
fx.rocdl.cvt_f32_fp8wrapper for the scalar ROCDL conversionMotivation
RDNA4 SplitKV needs scalar FP8 conversion for irregular key/value loads where a packed conversion is not applicable. The underlying MLIR ROCDL operation already exists, but FlyDSL does not currently expose it.
API stability
This is an additive stable API change. The stable API catalog gains
flydsl.expr.rocdl.cvt_f32_fp8; no existing entries are removed or changed.Validation
bash scripts/check_python_style.shpython3 -m pytest tests/unit/test_rocdl_conversions.py -qmain