[CK_TILE] Multiple-D GEMM example (#2219)

* Multiple d, initial commit * Check Ds Layout * Readme and clang format * Update branch & conflicts * Multiple D - fix clang-formatter * Rename elemetwise_op * Fix CI * Code review part1 * Remove printf * Remove unnecessary comment * Add new tests with Col layout * Review part 2 * Added support for Multiple D GEMM * Update comment * Remove maybe_unused * Clang-format * Review part 3 * Add comment to function * Add comment to function: another * Take number of params for a refrence function * Remove additional d param for 0 tensor * Change name of function * Fix CI fails
2026-05-02 12:41:26 +00:00 · 2025-06-13 19:39:11 +02:00
parent 3a0cb27966
commit bd96ac9742
34 changed files with 2267 additions and 285 deletions
--- a/example/ck_tile/19_gemm_multi_d/README.md
+++ b/example/ck_tile/19_gemm_multi_d/README.md
@@ -0,0 +1,35 @@
+#Multiple D GEMM
+
+This folder contains example for Multiple D GEMM using ck_tile tile-programming implementation.
+
+## build
+```
+#in the root of ck_tile
+mkdir build && cd build
+#you can replace < arch> with the appropriate architecture(for example gfx90a or gfx942) or \
+    leave it blank
+sh ../script/cmake-ck-dev.sh  ../ <arch>
+#The basic pipeline method on the gemm calculation
+make tile_example_gemm_multi_d_fp16 -j
+```
+This will result in an executable `build/bin/tile_example_gemm_multi_d_fp16`
+
+## example
+```
+args:
+       -m  M dimensions - (Default: 3840)
+       -n  N dimensions - (Default: 4096)
+       -k  K dimensions - (Default: 4096)
+-a_layout  Tensor A layout (default:R)
+-b_layout  Tensor B layout (default:C)
+-ds_layout Tensor D layout (default:R)
+-e_layout  Tensor E layout (default:R)
+-stride_a  Tensor A strides - (Default: 0)
+-stride_b  Tensor B strides - (Default: 0)
+-stride_e  Tensor C strides - (Default: 0)
+-stride_ds Tensor D strides - (Default: 0)
+-validate  0. No validation, 1. Validation on GPU. (Default: 1)
+  -warmup  Number of iterations before benchmark the kernel. (Default: 10)
+  -repeat  Number of iterations to benchmark the kernel. (Default: 100)
+  -kbatch  kbatch for SplitK. (Default 1)
+```