composable_kernel

mirror of https://github.com/ROCm/composable_kernel.git synced 2026-07-11 09:40:51 +00:00

Files

chris-tsiaousis-hpc 89c5e67028 [CK Tile] Unification work - mma transformations pipeline (#5508 )

## Motivation

In this PR we showcase how the amdgcn structs could be used in a pipeline that does some extra pre/post processing.
For the sparse intrinsics, so far we compressed the A vector "on the fly" right before the execution of the builtin. This might introduce performance issues down the line if, for example, the user decided to chain multiple sparse builtins. We tackle this problem by creating a specific SparseCompressTransform.

A MmaPipelineBase is also created to facilitate those kind of higher level compositions of the amdgcn structs and is integrated to the existing WaveWiseMma prototype. There is an effort to facilitate future operations, like swizzle A/B, C transpose or double/quad attr num access through the MmaPipelineOptionFlags, but those are not yet defined and should do so in a future PR.
The pipeline base class is basically at the RFC stage.

We also create a runtime test for the existing WaveWiseMma, as well as one for the SparseMma pipeline.

## Technical Details

The goal should be to have the pipeline easily expandable. May the CRTP of the base class or the interface in general be insufficient or unable to handle all of our needs, then a design modification should be discussed.

## Test Plan

New tests are added.

## Test Result

Tests should pass.

---------

Signed-off-by: Chris Tsiaousis <chris.tsiaousis@streamhpc.com>

2026-04-14 09:25:01 +02:00

arch

[CK Tile] Unification work - mma transformations pipeline (#5508 )

2026-04-14 09:25:01 +02:00

container

[CK_TILE] Optimize static_ford and sequence compile-time infrastructure (#5938 )

2026-04-02 15:25:14 -06:00

CMakeLists.txt

[CK TILE] Refactor sequence_reverse_inclusive_scan (#4355 )

2026-02-23 14:12:03 -07:00