Repository navigation
[LLVM] Fix true16 lowering for packed FP8 conversions - #1110
Open
big-yellow-duck wants to merge 3 commits into
Open
big-yellow-duck wants to merge 3 commits into
big-yellow-duck wants to merge 3 commits into
Conversation
Patch the pinned LLVM AMDGPU backend to select the real true16 packed FP8/BF8 conversion opcode and the low VGPR half on true16 targets. Apply the patch in local and CI LLVM builds and include all LLVM patch inputs in cache keys. Co-authored-by: vllmellm <190700713+vllmellm@users.noreply.github.com> Co-authored-by: Tan Pin Siang <1716735+tanpinsiang@users.noreply.github.com> Signed-off-by: big_yellow_duck <83417790+big-yellow-duck@users.noreply.github.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Motivation
llvm.amdgcn.cvt.pk.f32.{fp8,bf8}is selected through a fake16 pseudo because the intrinsic carries two selectable 16-bit words. On targets using real true16 instructions, the MC lowering must explicitly select the true16 opcode and spell a VGPR source as its low half. Without this, packed conversion lowering can produce an invalid or incorrectly encoded instruction.This PR is independent of the FlyDSL scalar
cvt_f32_fp8API addition and is based directly onmainfor focused review.Validation
git apply --checkagainst the pinned LLVM revisione2a39f504fee836e4def9581bed817ecc327b9dcllvm-lit -v llvm/test/CodeGen/AMDGPU/llvm.amdgcn.cvt.fp8.llpassesgit clang-format --diffreports no LLVM source formatting changes