Files
composable_kernel/include/ck_tile/core
Wojciech Laskowski c2601f38b7 [rocm-libraries] ROCm/rocm-libraries#6569 (commit 393049e)
Adding amdgcn_mma specializations for sparse MFMA builtins
 (#6569)

## Motivation

This PR is part of the [WMMA/MFMA] unification work. It's the fourth of
the series of PRs (after
https://github.com/ROCm/rocm-libraries/pull/5801,
https://github.com/ROCm/rocm-libraries/pull/6014 and
https://github.com/ROCm/rocm-libraries/pull/6567) that add all the
necessary MMA builtins as amdgcn_mma structs. This PR focuses on sparse
MFMA intrinsics.

## Technical Details

This change adds new specializations for MFMA sparse builtins. In total,
we add 27 MFMA builtins.

## Test Plan

All the new wrappers were added to the test suite in
`test_amdgcn_mma_layout.inc`.

## Test Result

Test pass locally, waiting for the CI.

## Submission Checklist

- [x] Look over the contributing guidelines at
https://github.com/ROCm/ROCm/blob/develop/CONTRIBUTING.md#pull-requests.
2026-06-12 12:48:29 +00:00
..
2024-04-15 19:27:12 -05:00

ck_tile/core

ck_tile/core contains every basic functions and structures to create a GPU kernel using ck_tile. User should only include ck_tile/core.hpp this single header to use all the functionality. Everything is under ck_tile namespace. The coding style under this folder should be similar to std (snake_case for structure/function, Camel for template types...)

algorithm/
    coordinate transform and some other reusable algorithm
arch/
    contains some basic device building block like mma, buffer addressing, etc...
container/
    contains basic container data structure, array/sequence/tuple/...
numeric/
    data type, and data type related math
tensor/
    tensor descriptors and tile level API
utility/
    other utility function for both host/device