mirror of https://github.com/ROCm/composable_kernel.git synced 2026-05-01 20:21:23 +00:00

Files

jakpiase 0bcb804ad0 [CK_TILE] Remove scratch usage from universal gemm (#2001 )

* moves kbatch condition outside of kernel

* add reviewer comments

* fixes

* fix tests

* fixes after review

---------

Co-authored-by: Adam Osewski <19374865+aosewski@users.noreply.github.com>

2025-05-05 18:46:44 +02:00

CMakeLists.txt

Revert "Add ck tile examples to package (#1880 )" (#2150 )

2025-04-30 10:20:16 -07:00

grouped_gemm.cpp

[CK_TILE] Remove scratch usage from universal gemm (#2001 )

2025-05-05 18:46:44 +02:00

grouped_gemm.hpp

[CK_TILE] Switch to universal gemm for batched and grouped gemms (#1919 )

2025-03-20 11:17:04 +01:00

README.md

Ck tile grouped GEMM example (#1713 )

2024-12-04 21:40:01 +01:00

run_grouped_gemm_example.inc

[CK_TILE] Switch to universal gemm for batched and grouped gemms (#1919 )

2025-03-20 11:17:04 +01:00

README.md

Grouped CShuffle GEMM

This folder contains example for Grouped GEMM using ck_tile tile-programming implementation. Currently, it only supports the basic feature of the CK Tile GEMM, but creates the placeholders for the future support on different GEMM pipeline and different GEMM modules. In the near future, we will gradually migrate all the GEMM features from old CK to CK Tile.

build

# in the root of ck_tile
mkdir build && cd build
# you can replace <arch> with the appropriate architecture (for example gfx90a or gfx942) or leave it blank
sh ../script/cmake-ck-dev.sh  ../ <arch>
# The basic pipeline method on the gemm calculation
make tile_example_grouped_gemm -j

This will result in an executable build/bin/tile_example_grouped_gemm

example

args:
   -a_layout    Tensor A layout (default:R)
   -b_layout    Tensor B layout (default:R)
   -c_layout    Tensor C layout (default:R)
          -v    0. No validation, 1. Validation on CPU
     -warmup    number of iterations before benchmark the kernel (default:10)
     -repeat    number of iterations to benchmark the kernel (default:100)