mirror of https://github.com/ROCm/composable_kernel.git synced 2026-07-12 02:05:50 +00:00

Files

Po Yen Chen f2dd2e5b09 Rangify constructor of HostTensorDescriptor & Tensor<> (#445 )

* Rangify STL algorithms

This commit adapts rangified std::copy(), std::fill() & std::transform()

* Rangify check_err()

By rangifying check_err(), we can not only compare values between
std::vector<>s, but also compare any ranges which have same value
type.

* Allow constructing Tensor<> like a HostTensorDescriptor

* Simplify Tensor<> object construction logics

* Remove more unnecessary 'HostTensorDescriptor' objects

* Re-format example code

* Re-write more HostTensorDescriptor ctor call

[ROCm/composable_kernel commit: 4a2a56c22f]

2022-11-11 11:36:01 -06:00

CMakeLists.txt

Batchnorm-forward and Batchnorm-infer Implemented using generic kernels (#320 )

2022-08-15 10:11:02 -05:00

dual_reduce_common.hpp

Rangify constructor of HostTensorDescriptor & Tensor<> (#445 )

2022-11-11 11:36:01 -06:00

dual_reduce_multiblock.cpp

Refactor device op implementations into impl subdirectory. (#420 )

2022-10-13 09:05:08 -05:00

dual_reduce_threadwise.cpp

Refactor device op implementations into impl subdirectory. (#420 )

2022-10-13 09:05:08 -05:00

README.md

Batchnorm-forward and Batchnorm-infer Implemented using generic kernels (#320 )

2022-08-15 10:11:02 -05:00

README.md

Instructions for `example_dual_reduce`

Run `example_dual_reduce_multiblock`

# -D <xxx> : input 4-d tensor lengths
# -v <x> :   verification (0=no, 1=yes)
#arg1: initialization (0=no init, 1=single integer value, 2=scope integer value, 3=decimal value)
#arg2: time kernel (0=no, 1=yes) 
./bin/example_dual_reduce_multiblock -D 600,28,28,256 -v 1 2 1

Result

./bin/example_dual_reduce_multiblock -D 600,28,28,256 -v 1 2 1                        
launch_and_time_kernel: grid_dim {150, 1, 1}, block_dim {256, 1, 1} 
Warm up 1 time
Start running 10 times...
Perf: 1.19529 ms, 201.499 GB/s, DeviceMultipleReduceBlockWise<256,M_C4_S1,K_C64_S1,InSrcVectorDim_1_InSrcVectorSize_1,OutDstVectorSize_1_1>

Run `example_dual_reduce_threadwise`

# -D <xxx> : input 4-d tensor lengths
# -v <x> :   verification (0=no, 1=yes)
#arg1: initialization (0=no init, 1=single integer value, 2=scope integer value, 3=decimal value)
#arg2: time kernel (0=no, 1=yes)
./bin/example_dual_reduce_multiblock -D 8000,4,4,4 -v 1 2 1

Result

./bin/example_dual_reduce_threadwise -D 8000,4,4,4 -v 1 2 1
launch_and_time_kernel: grid_dim {32, 1, 1}, block_dim {256, 1, 1} 
Warm up 1 time
Start running 10 times...
Perf: 0.01512 ms, 71.9577 GB/s, DeviceMultipleReduceThreadwise<256,M_C256_S1,K_C1_S4,InSrcVectorDim_1_InSrcVectorSize_2,OutDstVectorSize_1_1>

README.md

Instructions for example_dual_reduce

Run example_dual_reduce_multiblock

Run example_dual_reduce_threadwise

Instructions for `example_dual_reduce`

Run `example_dual_reduce_multiblock`

Run `example_dual_reduce_threadwise`