Split the instances by architecture. (#1223)

* parse examples inside the add_example_executable function * fix the example 64 cmake file * add xdl flag to the gemm_bias_softmax_gemm_permute example * add filtering of tests based on architecture type * enable test_grouped_gemm for gfx9 only * enable test_transpose only for gfx9 * only linnk test_transpose if it gets built * split the gemm instances by architectures * split gemm_bilinear,grouped_conv_bwd_weight instances by targets * split instances by architecture * split grouped_conv instances by architecture * fix clang format * fix the if-else logic in group_conv headers * small fix for grouped convolution instances * fix the grouped conv bwd weight dl instances * fix client examples * only enable client examples 3 and 4 on gfx9 * set the gfx9 macro * make sure the architecture macros are set by cmake * use separate set of xdl/wmma flags for host code * sinmplify the main cmake file * add conv_fwd_bf8 instance declaration [ROCm/composable_kernel commit: ae57e5938e]
2026-05-21 21:39:15 +00:00 · 2024-04-02 09:42:17 -07:00
parent 46ea205088
commit 1f4d13b2b5
160 changed files with 3770 additions and 3392 deletions
--- a/test/wrapper/CMakeLists.txt
+++ b/test/wrapper/CMakeLists.txt
@@ -12,10 +12,8 @@ add_dependencies(test_wrapper test_wrapper_copy)
 add_gtest_executable(test_wrapper_partition test_wrapper_partition.cpp)
 target_link_libraries(test_wrapper_partition PRIVATE utility)
 add_dependencies(test_wrapper test_wrapper_partition)
-if(GPU_TARGETS MATCHES "gfx908" OR GPU_TARGETS MATCHES "gfx90a" OR
-   GPU_TARGETS MATCHES "gfx940" OR GPU_TARGETS MATCHES "gfx941" OR
-   GPU_TARGETS MATCHES "gfx942")
-    add_gtest_executable(test_wrapper_gemm test_wrapper_gemm.cpp)
+add_gtest_executable(test_wrapper_gemm test_wrapper_gemm_xdl.cpp)
+if(result EQUAL 0)
    target_link_libraries(test_wrapper_gemm PRIVATE utility)
    add_dependencies(test_wrapper test_wrapper_gemm)
 endif()
--- a/test/wrapper/test_wrapper_gemm_xdl.cpp
+++ b/test/wrapper/test_wrapper_gemm_xdl.cpp