Chao Liu 66e4b458c4 Add gridwise GEMM pipeline (#89)
* clean up

* add mutilple thread scratch to ThreadwiseTensorSliceTransfer_v3r1

* add 2 stage prefetch

* add more sanity check into transform_tensor_descriptor

* tweak

* enabling 2 stage prefetch to exsiting gridwise gemm; tweak

* enabling 2 stage prefetch to exsiting gridwise gemm

* move gridwise gemm pipeline in class; clean up

* add some irregular tile size

* update CalculateHasMainK0BlockLoop for multi-stage-prefetch

* refactor gridwise gemm pipeline class

[ROCm/composable_kernel commit: 22d438ae9e]
2022-02-23 17:23:49 -06:00
2022-02-18 21:44:11 -06:00
2022-02-23 17:23:49 -06:00
2022-02-22 22:45:28 -06:00
2022-02-23 17:23:49 -06:00
2022-02-06 22:32:47 -06:00
2018-10-08 22:49:58 -05:00
2021-08-08 17:41:54 +00:00
2022-02-18 21:44:11 -06:00
2022-02-18 21:44:11 -06:00
2022-02-18 21:44:11 -06:00
2022-02-18 21:44:11 -06:00
2022-02-18 21:44:11 -06:00
Description
[DEPRECATED] Moved to ROCm/rocm-libraries repo. NOTE: develop branch is maintained as a read-only mirror
Readme MIT 234 MiB
Languages
C++ 93.1%
Python 4.5%
CMake 1.5%
Shell 0.5%
Pawn 0.2%