ROCm/composable_kernel

mirror of https://github.com/ROCm/composable_kernel.git synced 2026-05-14 02:02:46 +00:00

Files

History

ArthurLiu c1d2cf0869 [CK][CK_TILE] Fix FMHA codegen group mode dispatch (#6764 )

## Motivation

FMHA codegen had incorrect dispatch behavior in group mode. Two root
causes:

1. Wrong field names in dispatch conditions — Used batch-mode fields
(seqlen_q, seqlen_k) instead of group-mode fields (max_seqlen_q,
max_seqlen_k), causing wrong kernel selection at runtime on gfx950.
2. Missing kernel variants — Group mode was overly filtered out from
smaller-tile specializations (bwd) and lacked spatial-padding pipeline
variants on gfx950 (fwd).

gfx942 don't support trload pipeline. 

## Technical Details

 fmha_bwd.py:
- max_seq_q_cond and extra_cond now emit t.max_seqlen_q / t.max_seqlen_k
for group mode.
- Relaxed kernel filtering: group mode no longer skips tiles with
max_seq_q != 0.

  fmha_fwd.py:
  - get_bm0_cond emits a.max_seqlen_q for group mode tile-size dispatch.
- Added two qr_async_trload pipeline variants with spatial padding for
gfx950 group mode.
  
## Test Plan
Triggering AITER CI job:

## Submission Checklist

- [ x] Look over the contributing guidelines at
https://github.com/ROCm/ROCm/blob/develop/CONTRIBUTING.md#pull-requests.

2026-04-28 02:14:42 +08:00

..

Padding support for wave transfer (#3537 )

2026-01-26 12:57:09 -08:00

02_gemm_bilinear

[CI, CK examples] Disable time_kernel for CI tests and examples (#3464 )

2026-01-07 16:30:57 +01:00

03_gemm_bias_relu

chore(copyright) update library wide CMakeLists.txt copyright header template (#3313 )

2025-11-28 13:49:54 -08:00

04_gemm_add_add_fastgelu

chore(copyright) update library wide CMakeLists.txt copyright header template (#3313 )

2025-11-28 13:49:54 -08:00

[CK] Integrate GPU reference into ckProfiler for convolutions (#3379 )

2025-12-18 07:59:45 +01:00

10_convnd_fwd_multiple_d_multiple_reduce

chore(copyright) update library wide CMakeLists.txt copyright header template (#3313 )

2025-11-28 13:49:54 -08:00

11_convnd_fwd_bias

[DOCS] Documentation Addition (Readme updates) (#2495 )

2025-10-16 03:10:57 -07:00

[CK] Replace tuple value construction with tuple_element_t type extraction [1A] (#5030 )

2026-03-06 09:27:27 -07:00

[CI, CK examples] Disable time_kernel for CI tests and examples (#3464 )

2026-01-07 16:30:57 +01:00

14_gemm_quantization

[CK][Examples] Adding parameters for a couple of CK examples:

2026-03-12 09:47:41 +01:00

15_grouped_gemm

[CK] Fix/suppress clang lifetimebound warnings with staging compiler. (#6550 )

2026-04-22 15:47:47 +00:00

16_gemm_multi_d_multi_reduces

[CK][Examples] Adding parameters for a couple of CK examples:

2026-03-12 09:47:41 +01:00

17_convnd_bwd_data

[CK] Integrate GPU reference into ckProfiler for convolutions (#3379 )

2025-12-18 07:59:45 +01:00

18_batched_gemm_reduce

chore(copyright) update library wide CMakeLists.txt copyright header template (#3313 )

2025-11-28 13:49:54 -08:00

19_binary_elementwise

chore(copyright) update library wide CMakeLists.txt copyright header template (#3313 )

2025-11-28 13:49:54 -08:00

20_grouped_conv_bwd_weight

[CK] Small improvements for grouped conv backward weight (#4872 )

2026-02-25 20:10:12 +00:00

21_gemm_layernorm

chore(copyright) update library wide CMakeLists.txt copyright header template (#3313 )

2025-11-28 13:49:54 -08:00

[CI, CK examples] Disable time_kernel for CI tests and examples (#3464 )

2026-01-07 16:30:57 +01:00

[CI, CK examples] Disable time_kernel for CI tests and examples (#3464 )

2026-01-07 16:30:57 +01:00

24_batched_gemm

[CI, CK examples] Disable time_kernel for CI tests and examples (#3464 )

2026-01-07 16:30:57 +01:00

25_gemm_bias_e_permute

Implement batched gemm bias permute for RDNA4 (#3534 )

2026-01-17 08:30:27 +01:00

CK: Extract shared boilerplate from 47 gemm_quant test files (#6323 )

2026-04-11 06:00:26 -04:00

27_layernorm2d_fwd

chore(copyright) update library wide CMakeLists.txt copyright header template (#3313 )

2025-11-28 13:49:54 -08:00

28_grouped_gemm_bias_e_permute

[CK] Fix/suppress clang lifetimebound warnings with staging compiler. (#6550 )

2026-04-22 15:47:47 +00:00

29_batched_gemm_bias_e_permute

[CK] Fix/suppress clang lifetimebound warnings with staging compiler. (#6550 )

2026-04-22 15:47:47 +00:00

30_grouped_conv_fwd_multiple_d

[CI, CK examples] Disable time_kernel for CI tests and examples (#3464 )

2026-01-07 16:30:57 +01:00

31_batched_gemm_gemm

chore(copyright) update library wide CMakeLists.txt copyright header template (#3313 )

2025-11-28 13:49:54 -08:00

32_batched_gemm_scale_softmax_gemm

chore(copyright) update library wide CMakeLists.txt copyright header template (#3313 )

2025-11-28 13:49:54 -08:00

33_multiple_reduce

[CI, CK examples] Disable time_kernel for CI tests and examples (#3464 )

2026-01-07 16:30:57 +01:00

chore(copyright) update library wide CMakeLists.txt copyright header template (#3313 )

2025-11-28 13:49:54 -08:00

Revert "[ck] Support VGPR estimate in GridwiseGemm_wmma_cshuffle_v3" (#4762 )

2026-02-20 22:40:28 +00:00

36_sparse_embedding

[CK] Fix/suppress clang lifetimebound warnings with staging compiler. (#6550 )

2026-04-22 15:47:47 +00:00

37_batched_gemm_add_add_relu_gemm_add

Implement batched gemm add relu gemm add for rdna4 (#3391 )

2026-01-20 13:06:59 -08:00

38_grouped_conv_bwd_data_multiple_d

Grouped convolution backward data WMMA v3 implementation (#3460 )

2025-12-30 16:25:08 +01:00

[CK] Fix/suppress clang lifetimebound warnings with staging compiler. (#6550 )

2026-04-22 15:47:47 +00:00

40_conv2d_fwd_quantization

Fix per-layer conv2d int8 CPU verification reference path (#6656 )

2026-04-23 07:08:50 -07:00

41_grouped_conv_conv_fwd

chore(copyright) update library wide CMakeLists.txt copyright header template (#3313 )

2025-11-28 13:49:54 -08:00

42_groupnorm_fwd

[CI, CK examples] Disable time_kernel for CI tests and examples (#3464 )

2026-01-07 16:30:57 +01:00

43_splitk_gemm_bias_e_permute

[CK] Fix/suppress clang lifetimebound warnings with staging compiler. (#6550 )

2026-04-22 15:47:47 +00:00

44_elementwise_permute

[CI, CK examples] Disable time_kernel for CI tests and examples (#3464 )

2026-01-07 16:30:57 +01:00

45_elementwise_normalization

[CI, CK examples] Disable time_kernel for CI tests and examples (#3464 )

2026-01-07 16:30:57 +01:00

46_gemm_add_multiply

chore(copyright) update library wide CMakeLists.txt copyright header template (#3313 )

2025-11-28 13:49:54 -08:00

47_gemm_bias_softmax_gemm_permute

chore(copyright) update library wide CMakeLists.txt copyright header template (#3313 )

2025-11-28 13:49:54 -08:00

chore(copyright) update library wide CMakeLists.txt copyright header template (#3313 )

2025-11-28 13:49:54 -08:00

49_maxpool2d_bwd

chore(copyright) update library wide CMakeLists.txt copyright header template (#3313 )

2025-11-28 13:49:54 -08:00

chore(copyright) update library wide CMakeLists.txt copyright header template (#3313 )

2025-11-28 13:49:54 -08:00

51_avgpool3d_bwd

chore(copyright) update library wide CMakeLists.txt copyright header template (#3313 )

2025-11-28 13:49:54 -08:00

52_im2col_col2im

chore(copyright) update library wide CMakeLists.txt copyright header template (#3313 )

2025-11-28 13:49:54 -08:00

53_layernorm2d_bwd

chore(copyright) update library wide CMakeLists.txt copyright header template (#3313 )

2025-11-28 13:49:54 -08:00

54_groupnorm_bwd

chore(copyright) update library wide CMakeLists.txt copyright header template (#3313 )

2025-11-28 13:49:54 -08:00

59_grouped_gemm_multi_ABD

[CK] Implement device grouped gemm fixed nk multi abd for rdna4 (#4425 )

2026-02-25 05:16:07 +00:00

60_gemm_multi_ABD

CK: Remove 41 commented-out dead code blocks (~200 lines) (#6302 )

2026-04-10 11:17:11 -04:00

61_contraction_multi_ABD

chore(copyright) update library wide CMakeLists.txt copyright header template (#3313 )

2025-11-28 13:49:54 -08:00

62_convnd_activ

Adding remaining conv, dynamic_op, and scaleadd_scaleadd_relu flavors for grouped conv fwd (#3529 )

2026-01-30 17:02:14 +01:00

63_layernorm4d_fwd

chore(copyright) update library wide CMakeLists.txt copyright header template (#3313 )

2025-11-28 13:49:54 -08:00

64_fpAintB_gemm

chore(copyright) update library wide CMakeLists.txt copyright header template (#3313 )

2025-11-28 13:49:54 -08:00

65_gemm_multiply_multiply

CK: Remove 41 commented-out dead code blocks (~200 lines) (#6302 )

2026-04-10 11:17:11 -04:00

66_complex_contraction_bilinear

CK: Extract shared boilerplate from 47 gemm_quant test files (#6323 )

2026-04-11 06:00:26 -04:00

67_gemm_microscaling

[CI, CK examples] Disable time_kernel for CI tests and examples (#3464 )

2026-01-07 16:30:57 +01:00

[CK][Examples] Fixing stride issues in ck examples 14/65/68/69 by workaround - Bypassing hostTensor validation

2026-01-15 16:43:02 +01:00

69_gemm_add_relu

[CK][Examples] Fixing stride issues in ck examples 14/65/68/69 by workaround - Bypassing hostTensor validation

2026-01-15 16:43:02 +01:00

[CK][CK_TILE] Fix FMHA codegen group mode dispatch (#6764 )

2026-04-28 02:14:42 +08:00

CMakeLists.txt

Build CK on Windows (#3458 )

2026-01-14 07:31:45 -08:00

README.md

Add basic documentation structure (#1715 )

2024-12-04 00:46:47 +01:00

README.md

Back to the main page

Composable Kernel examples