composable_kernel

mirror of https://github.com/ROCm/composable_kernel.git synced 2026-05-11 00:40:09 +00:00

Files

ltqin 10b3278b05 Skip lds of b matrix (#326 )

* start

* read for gridwise gemm

* add MakeBGridDescriptor_K0_N0_N1_N2_N3_K1

* add thread  copy desc and register buffer

* add K0PerBlock dim

* add read global data

* finish gridwise gemm

* finish blockwise gemm

* add print data

* add smallest config

* add compare code for gridwis gemm

* fix NXdlPerWave

* fix k0perthread and gridewis gemm main loop

* remove b matrix lds alloc

* fix name

* add test code

* create b_grid_desc_k0_k1_k2_n0_n1_n2_n3_k3 from parameter

* add double register

* modify b_thread_desc_

* add float

* fp16 tag

* add tail for pipeline

* finish main loop

* optimize main loop

* start clear gridwise gemm

* clear code

* clear redundant code

* change file name

* change file name

* fix bug after merge develop

* fix input parameters

* using MultiK0 control b load data loop

* fix some config

* 4 buffer

* fix bug

* one can use

* change read order

* change buffer array to tuple

* change to 8 buffer

* interleave buffer load

* change to 16

* read 8 buffer

* add data buffer to template

* fix after merge develop(head file)

* format

* change to 4 buffer

* remove unnecessary lambda fun

2022-08-13 01:35:49 -05:00

reduction_functions_threadwise.hpp

Single-kernel GEMM + layernorm (#263 )

2022-07-01 01:38:00 -05:00

threadwise_contraction_dl.hpp

add license in file (#303 )

2022-06-24 23:32:43 -05:00

threadwise_gemm_dlops_v3.hpp

add license in file (#303 )

2022-06-24 23:32:43 -05:00

threadwise_tensor_slice_set.hpp

add license in file (#303 )

2022-06-24 23:32:43 -05:00

threadwise_tensor_slice_transfer_v3r1.hpp

add license in file (#303 )

2022-06-24 23:32:43 -05:00

threadwise_tensor_slice_transfer_v3r3.hpp

add license in file (#303 )

2022-06-24 23:32:43 -05:00

threadwise_tensor_slice_transfer_v4r1.hpp

add license in file (#303 )

2022-06-24 23:32:43 -05:00

threadwise_tensor_slice_transfer_v5r1.hpp

add license in file (#303 )

2022-06-24 23:32:43 -05:00

threadwise_tensor_slice_transfer_v6r1.hpp

add license in file (#303 )

2022-06-24 23:32:43 -05:00

threadwise_tensor_slice_transfer_v6r2.hpp

add license in file (#303 )

2022-06-24 23:32:43 -05:00

threadwise_tensor_slice_transfer_v6r3.hpp

add license in file (#303 )

2022-06-24 23:32:43 -05:00

threadwise_tensor_slice_transfer_v7.hpp

add license in file (#303 )

2022-06-24 23:32:43 -05:00

threadwise_tensor_slice_transfer.hpp

Skip lds of b matrix (#326 )

2022-08-13 01:35:49 -05:00