Repository navigation
[Kernel][Perf] Optimize gfx950 head64 prefill attention - #1126
Open
michael604work wants to merge 1 commit into
Open
michael604work wants to merge 1 commit into
michael604work wants to merge 1 commit into
Conversation
Add specialized dense, paged, packed-varlen, global, and sliding-window paths for BF16 GQA with head dimension 64 and page size 64. Improve causal scheduling, page-table address handling, and the dual-wave softmax/MFMA pipeline for long-context workloads.
Collaborator
|
Conflict. and perf compare with other hdim and existing solution? |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Optimize gfx950 BF16 prefill attention for the 42B MXFP4 serving shape:
The change adds workload-specific routing and improves the dual-wave/generic
pipelines through causal work reordering, mask specialization, page-table and
page-address reuse, overlapped K/V address preparation, and softmax/MFMA
scheduling.
Motivation
The existing gfx950 attention paths did not cover this head-dim-64/page-64
serving workload with the required combination of long-context paged KV,
packed varlen requests, and sliding-window attention. This is needed for the
42B MXFP4 model configuration.
Workload
gfx950)Performance
Representative dense-global throughput for the accepted campaign incumbent,
using causal-triangle QK+PV FLOPs and the MI350X 2.3 PFLOP/s BF16 peak:
Each campaign promoted changes only after same-GPU paired measurements showed
at least 0.5% incremental geometric-mean improvement, no workload regression
above 3%, candidate MAD at most 3%, and all correctness gates passing.
Test plan
scripts/check_python_style.sh --include-localpytest tests/kernels/test_swa_gfx950.py -q(5 passed)tables, including sampled-reference comparison
tables, including sampled-reference comparison
including non-page-aligned lengths and KV up to 128K; zero mismatches
mismatches
Kernel-level tests and microbenchmarks only; no model/e2e tests were run.