mirror of https://github.com/ROCm/composable_kernel.git synced 2026-07-17 09:08:35 +00:00

Go to file

rocking5566 88e9bfd4da Standalone layernorm (#315 )

* Implement layernorm kernel and deviceOp

* verify gpu kernel with host code

* 1. Separate gamma aand beta from affine
2. Check if argument is valid

* clean

* Sync the naming

* Support sweep once mode if we can put k dimension data inside one block

* [What] Get length from upper length.
[Why] if we get length directly, we may get length after padding.

* We only use one block in K dimension.
Hence, we can simplify the indexing of global R/W.

* Use 1d descriptor for gamma and beta

* Add accElementwiseOp

* Extract layernorm host code

* Support different YVectorDim in GridwiseLayernorm

* Rename XSrcVectorDim to XYSrcVectorDim. Because we use same parameter in deviceOp

* Gamma and beta can share the VGPR.

* Add test for fp32 and fp16

* Fix bug of concurrency and add test case which may fail orignally

* Propagate NaN for layernorm

Co-authored-by: Chao Liu <chao.liu2@amd.com>

[ROCm/composable_kernel commit: 7f21662089]

2022-07-13 11:16:14 -05:00

client_example

minor fix in gemm client example (#328 )

2022-07-13 10:54:38 -05:00

cmake

Switch to standard ROCm packaging (#301 )

2022-06-25 09:35:16 -05:00

example

Standalone layernorm (#315 )

2022-07-13 11:16:14 -05:00

include/ck

Standalone layernorm (#315 )

2022-07-13 11:16:14 -05:00

library

Standalone layernorm (#315 )

2022-07-13 11:16:14 -05:00

profiler

add conv1d/3d bwd weight instances (#318 )

2022-07-08 15:42:20 -05:00

script

Add switch between compilers, make 9110 compiler default, add full QA scripts. (#322 )

2022-07-13 09:27:43 -05:00

test

Standalone layernorm (#315 )

2022-07-13 11:16:14 -05:00

.clang-format

start adding convolution

2018-10-08 22:49:58 -05:00

.clang-tidy

add tidy

2021-08-08 17:41:54 +00:00

.gitignore

Switch to standard ROCm packaging (#301 )

2022-06-25 09:35:16 -05:00

CMakeLists.txt

Remove incorrect old packaging statement (#308 )

2022-06-30 09:40:03 -05:00

Config.cmake.in

Add host API (#220 )

2022-05-12 09:21:01 -05:00

dev-requirements.txt

Initial Setup for CI (#86 )

2022-02-18 21:44:11 -06:00

Dockerfile

Add switch between compilers, make 9110 compiler default, add full QA scripts. (#322 )

2022-07-13 09:27:43 -05:00

Jenkinsfile

Add switch between compilers, make 9110 compiler default, add full QA scripts. (#322 )

2022-07-13 09:27:43 -05:00

LICENSE

update license (#297 )

2022-06-23 01:27:30 -05:00

rbuild.ini

Update test CMakeLists to add new tests automatically and add Jenkins stage for tests (#88 )

2022-03-03 16:59:42 -06:00

README.md

Improve external interface for GEMM and GEMM+add+add+fastgelu (#311 )

2022-06-30 22:11:00 -05:00

requirements.txt

Update test CMakeLists to add new tests automatically and add Jenkins stage for tests (#88 )

2022-03-03 16:59:42 -06:00

README.md

Docker script

docker run                                     \
-it                                            \
--privileged                                   \
--group-add sudo                               \
-w /root/workspace                             \
-v ${PATH_TO_LOCAL_WORKSPACE}:/root/workspace  \
rocm/tensorflow:rocm5.1-tf2.6-dev              \
/bin/bash

Install the new rocm-cmake version

https://github.com/RadeonOpenCompute/rocm-cmake

Build

mkdir build && cd build

# Need to specify target ID, example below is gfx908 and gfx90a
cmake                                                                 \
-D BUILD_DEV=OFF                                                      \
-D CMAKE_BUILD_TYPE=Release                                           \
-D CMAKE_CXX_FLAGS=" --offload-arch=gfx908 --offload-arch=gfx90a -O3" \
-D CMAKE_CXX_COMPILER=/opt/rocm/bin/hipcc                             \
-D CMAKE_PREFIX_PATH=/opt/rocm                                        \
-D CMAKE_INSTALL_PREFIX=${PATH_TO_CK_INSTALL_DIRECTORY}               \
..

Build and Run Examples

 make -j examples

Instructions for running each individual examples are under example/

Tests

 make -j examples tests
 make test

Build ckProfiler

 make -j ckProfiler

Instructions for running ckProfiler are under profiler/

Install CK

make install

Using CK as pre-built kernel library

Caveat

Kernel Timing and Verification

CK's own kernel timer will warn up kernel once, and then run it multiple times to get average kernel time. For some kernels that use atomic add, this will cause output buffer to be accumulated multiple times, causing verfication failure. To work around it, do not use CK's own timer and do verification at the same time. CK's own timer and verification in each example and ckProfiler can be enabled or disabled from command line.

Languages

C++ 90.7%

Python 6.6%

CMake 1.7%

Shell 0.5%

Pawn 0.2%

Other 0.1%