composable_kernel

mirror of https://github.com/ROCm/composable_kernel.git synced 2026-07-19 02:01:01 +00:00

Author	SHA1	Message	Date
trixirt	efaf31061a	cmake: Add CK_PARALLEL_LINK_JOBS and CK_PARALLEL_COMPILE_JOBS options (#1063 ) Copied from the llvm-project LLVM_PARALLEL_*_JOBS Concurrent linking can break the build as well as having too many compile jobs for the avaiable memory. These options allow the user to fine tune the build to fit within their machines memory constraints. An example use on linux is COMPILE_JOBS=`cat /proc/cpuinfo \| grep -m 1 'cpu cores' \| awk '{ print $4 }'` if [ ${COMPILE_JOBS}x = x ]; then COMPILE_JOBS=1 fi BUILD_MEM=4 MEM_KB=0 MEM_KB=`cat /proc/meminfo \| grep MemTotal \| awk '{ print $2 }'` MEM_MB=`eval "expr ${MEM_KB} / 1024"` MEM_GB=`eval "expr ${MEM_MB} / 1024"` COMPILE_JOBS_MEM=`eval "expr 1 + ${MEM_GB} / ${BUILD_MEM}"` if [ "$COMPILE_JOBS_MEM" -lt "$COMPILE_JOBS" ]; then COMPILE_JOBS=$COMPILE_JOBS_MEM fi LINK_MEM=32 LINK_JOBS=`eval "expr 1 + ${MEM_GB} / ${LINK_MEM}"` cmake -G Ninja -DCK_PARALLEL_LINK_JOBS=$LINK_JOBS -DCK_PARALLEL_COMPILE_JOBS=$COMPILE_JOBS Signed-off-by: Tom Rix <trix@redhat.com>	2023-12-14 17:26:41 -08:00
Lisa	281f836903	fix typo (#1067 ) Co-authored-by: Illia Silin <98187287+illsilin@users.noreply.github.com>	2023-12-14 14:21:18 -08:00
Jun Liu	3a3b98ef79	[Doc][Werror] Fix security alerts and sync with MIOpen (#1085 ) * fix Werror unused-parameter * sync doc requirements * fix blank space format * fix dependency issue	2023-12-13 12:50:15 -08:00
Rostyslav Geyyer	6891e4d109	Fix the bugs (#1099 )	2023-12-13 12:27:31 -08:00
Illia Silin	c004e0d990	disabling some fp8 gemm instances to reduce build time (#1084 ) * disabling some fp8 gemm instances to reduce build time * disable fp8 gemm instances to reduce build time * remove the unused variable * build fp8 gemm default and padded instances separately * fix include pathsc	2023-12-11 17:49:27 -08:00
Bartlomiej Wroblewski	89ee47460b	Fix IsSupported check in the contraction op (#1066 ) Current implementation of IsSupported method in contraction ops does not cover a lot of possible cases in which ScalarPerVector cannot really be used to read A, B or D, or write E. This PR extends both the regular and multiABD contraction ops with improved checks and also adds new instances with smaller values of ScalarPerVector to support instances that are not supported by other instances.	2023-12-11 17:12:32 +01:00
Illia Silin	f199035b74	fix clang format (#1095 )	2023-12-08 14:32:37 -08:00
Nicolas Macchioni	b4dcd5803f	Add F8 dtype definition in f16_f8_f16 gemm instances (#1092 )	2023-12-08 13:30:01 -06:00
Bartłomiej Kocot	f836984891	Support broadcast for bias in grouped conv fwd (#1081 ) * Support broadcast for bias in grouped conv fwd * Fix comment * Comment fixes * Remove GK layout	2023-12-08 11:07:42 +01:00
Illia Silin	d939411dae	Switch from ROCmSoftwarePlatform to ROCm org (#1091 ) * switch from ROCmSoftwarePlatform to ROCm org * replace ROCmSoftwarePlatform with ROCm in few more places	2023-12-07 15:59:34 -08:00
zjing14	33600202c6	remove imcomplete transpose profiler (#1088 ) Co-authored-by: Jing Zhang <jizha@amd.com> Co-authored-by: Illia Silin <98187287+illsilin@users.noreply.github.com>	2023-12-07 13:39:40 -06:00
dependabot[bot]	957281ce45	Bump rocm-docs-core from 0.29.0 to 0.30.1 in /docs/sphinx (#1090 ) Bumps [rocm-docs-core](https://github.com/RadeonOpenCompute/rocm-docs-core) from 0.29.0 to 0.30.1. - [Release notes](https://github.com/RadeonOpenCompute/rocm-docs-core/releases) - [Changelog](https://github.com/RadeonOpenCompute/rocm-docs-core/blob/develop/CHANGELOG.md) - [Commits](https://github.com/RadeonOpenCompute/rocm-docs-core/compare/v0.29.0...v0.30.1) --- updated-dependencies: - dependency-name: rocm-docs-core dependency-type: direct:production update-type: version-update:semver-minor ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>	2023-12-07 10:32:04 -07:00
Illia Silin	6896c3b0ae	Fix the CI builds using clang++ directly. (#1087 ) * turn on -O3 compiler flag explicitly * change cmake syntax for CI * modify cmake line breaks in jenkinsfile	2023-12-06 12:48:10 -08:00
Bartłomiej Kocot	836b7e557d	Introduce wrapper library (#1071 ) * Introduce wrapper library * Update cmake files * Revert "Update cmake files" This reverts commit `c27f88b565`. * Fix comments	2023-12-06 11:58:59 +01:00
Sam Wu	f60cd9d7a6	Standardize documentation for ReadtheDocs (#1057 ) Relates to https://github.com/RadeonOpenCompute/rocm-docs-core/issues/330	2023-12-05 11:05:55 -07:00
Jun Liu	ff24b537cb	[SWDEV-435347] disable instances failed with mainlien compiler (#1077 )	2023-12-04 23:45:16 -08:00
Illia Silin	afe4622014	Add daily run with mainline compiler. (#1075 ) * add daily build with mainline compiler * fix the compiler paths for ci * remove the -flto flag * build with clang by default	2023-12-04 19:04:52 -08:00
Bartlomiej Wroblewski	bc4bf9bd03	Add support for double buffering in direct load GEMM kernel (#1052 ) This PR introduces support for double buffering in LDS into GEMM kernels that use direct load instructions. Direct loads now use inline asm instead of intrinsics. Usage of intrinsics results in compiler adding additional waitcnt instructions what breaks possible load/compute overlap in case of double buffering. Usage of inline asm results in the need to use sched_barrier in order to make sure that compiler cannot incorrectly reschedule instructions since it does not know the data dependencies between global->LDS and LDS->registers.	2023-12-03 23:08:47 +01:00
Jun Liu	c7d5c7727b	[CI] Update Jenkinsfile (#1073 )	2023-11-30 15:24:59 -08:00
zjing14	49df1dc595	Fixed GroupedGemmFixedNK with hipGraph (#1065 ) * fixed examples; add async_mem_set * add stream to all deviceOp using SetWorkspace --------- Co-authored-by: Jing Zhang <jizha@amd.com>	2023-11-30 15:09:27 -06:00
Bartłomiej Kocot	8ff845f2c4	Introduce wrapper for layout (#1054 ) * Introduce wrapper for layout * Extend functionality * Fix for getLength * Comment fixes * Add comments and remove not needed getters	2023-11-30 12:11:43 +01:00
arai713	a2969aa8b6	Disable transpose device op for MI300 (#1050 ) * added working example for 5D input using 1D kernel * example with 5D input tensor and 2d kernel - not working: issues with arguments * added updated version of 3d device op - changed descriptors/dims * added example file to check kernel * fixed descriptor and isSupportedArgument stride problem * added and modified kernel for 3d - updated tids/loop * adding some more 5d example files * fixed some issues * changes made for testing * working version: fixed error in stride for A, still a bit inefficient * cleaned up formatting/comments * updating formatting * more formatting fixes * fixing cmake, adding back gpu targets in cmake script * adding client example * added instances for client example * fixed errors in client example * implemented client ex with device_elementwise.hpp and device_elementwise_3d_impl.hpp * removed extra files * minor formatting and naming fixes * adding test files and profiler * fixing minor error * minor fix * removed unneccesary comments, renamed files * updated instance list for client example, added different layout example * removing instances * fixed error in instance generation * remove comments * update profiler and client example tensor layouts * fixed errors in test/profiler * updated vector dim access to enable vector load * updated test/profiler files * updated example with 1d kernel * updating profiler * renamed files * disabled device op for MI300 * skip elementwise_permute_2d on gfx94x * Update CMakeLists.txt * fixing CMake - disabling some GPU targets --------- Co-authored-by: Jing Zhang <jizha@amd.com> Co-authored-by: Jing Zhang <jizhan@amd.com> Co-authored-by: zjing14 <zhangjing14@gmail.com>	2023-11-29 11:36:40 -06:00
zjing14	ae5e5181aa	recover default niter (#1064 )	2023-11-28 12:18:42 -08:00
Illia Silin	7965d66a81	Split the static library into several files. (#1044 ) * spolit the static library into several * update lib paths and fix client example * do not use device_mha_operarions for client examples * use appropriate libs to link to client examples * remove the gpu/transpose path from the list * try fixing clinet examples 3,4,9 * add necessary libs for client examples * fix the layernorm client example * fix the client examples 23 and 24 * fix typo * add interface library and refresh clang format	2023-11-28 11:17:37 -08:00
Rostyslav Geyyer	6ef034f6ca	Switch default f8 conversion to stochastic rounding (#1048 ) * Switch default f8 conversion to stochastic rounding * Refactor f8-related type_converts * Add an element-wise op	2023-11-27 20:06:17 -06:00
Bartlomiej Wroblewski	60ecfd73f9	Add missing check for K padding in XDL GEMM (#1056 )	2023-11-27 11:31:39 +01:00
Bartlomiej Wroblewski	bfecc19352	Fix cluster length arrange order in fp16 GEMM example (#1055 )	2023-11-27 11:31:14 +01:00
Bartlomiej Wroblewski	627054b941	Add basic support for direct loads from global to LDS (#999 ) * Add basic support for direct loads from global to LDS * Clean the code and comments * Add support for fp16 * Add comments * Add check for thread cluster lengths * Align non-direct-load fp16 example * Small fixes * Extend IsSupported to check for supported GPU gens * Build examples only on the supported HW * Do not throw when instance not supported in 04 example * Review: Apply review suggestions * Review: small fix * Review: small fix	2023-11-25 13:35:22 +01:00
zjing14	e8cddfdc3b	Improve 4k gemm perf (#1047 ) * improve 4k gemm perf * add f8 instances * format --------- Co-authored-by: Jing Zhang <jizha@amd.com>	2023-11-17 07:06:24 -06:00
Chao Liu	e1fa00917c	[Hotfix] Remove unsed profile_transpose.cpp (#1046 )	2023-11-16 14:49:46 -08:00
dependabot[bot]	61cce232c7	Bump rocm-docs-core from 0.26.0 to 0.27.0 in /docs/sphinx (#1023 ) Bumps [rocm-docs-core](https://github.com/RadeonOpenCompute/rocm-docs-core) from 0.26.0 to 0.27.0. - [Release notes](https://github.com/RadeonOpenCompute/rocm-docs-core/releases) - [Changelog](https://github.com/RadeonOpenCompute/rocm-docs-core/blob/develop/CHANGELOG.md) - [Commits](https://github.com/RadeonOpenCompute/rocm-docs-core/compare/v0.26.0...v0.27.0) --- updated-dependencies: - dependency-name: rocm-docs-core dependency-type: direct:production update-type: version-update:semver-minor ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>	2023-11-15 22:32:23 -08:00
Bartłomiej Kocot	1fefd82ed8	Log CDEBlockTransferScalarPerVector_NPerBlock in conv fwd multiD xdl (#1042 ) * Log CDEBlockTransferScalarPerVector_NPerBlock in conv_fwd_multi_d_xdl implementation * Log CDEBlockTransferScalarPerVector_NPerBlock in conv fwd multiD xdl	2023-11-15 17:31:50 +01:00
Bartłomiej Kocot	3ef3102fc5	Fix check for conv Fwd Filter1x1Pad0 (#1040 ) * Fix check for conv Fwd Filter1x1Pad0 * Fix check for conv Fwd Filter1x1Pad0	2023-11-15 17:28:33 +01:00
Bartłomiej Kocot	f2398f612d	Introduce multiABD api and deprecate multiD (#1035 ) * Introduce multiABD api and deprecate multiD * Replace multiD with multiABD * Mark structures as deprecated * Change doxygen deprecated to note to avoid warnings	2023-11-14 17:00:40 +01:00
Rostyslav Geyyer	5356c4a943	Add conv bwd weight client example (#1005 ) * Add conv bwd weight client example * Update instance selector * Fake the conversion * Bring the conversion back	2023-11-13 11:16:04 -06:00
arai713	454cf7bd1f	Hip tensor permute (#1002 ) * adding files for F32 example * adding functioning implementation with scalar multiplication and unary operator support * added fp 16 type check in unary square * updating scalar multiplication as an operator * functioning version with scalar operator * changing strides for col major * updated column major implementation * working column major implementation * cleaned up comments, rearranged/renamed files	2023-11-13 11:15:48 -06:00
zjing14	600fc000ed	add more instances for bfp16 gemm (#1036 ) * add more instances for bfp16 * reduce the gemm input values to prevent round-off errors --------- Co-authored-by: Jing Zhang <jizha@amd.com> Co-authored-by: illsilin <Illia.Silin@amd.com>	2023-11-11 07:09:32 -08:00
Bartłomiej Kocot	49e52bb357	Support multi AB for grouped conv fwd xdl (#1027 ) * Support multi AB for grouped conv fwd xdl * Add instances * Add client example * Add example * Add interface test * Minor fixes Minor fixes Minor fixes * Comment fixes * Fixes * Reference fix * Test xdl fixes * Improve multi_ab interface test	2023-11-10 15:54:44 +01:00
rocking	1db7560365	Backward of gamma and beta for layernorm and groupnorm (#1013 ) * Add layernorm backward reference code * Add groupnorm backward reference code * Add example * clang format * Fixc bug of reference layernorm and groupnorm * Fix naming * Refine naming * Add device op for normalization bwd gamma and beta * Refine template parameter * Add bwd gamma & beta of kernel * 1. Add groupnorm example 2. Refine layernorm naming * Narrow down the static check for performance * Refine variable name	2023-11-10 18:02:03 +08:00
Illia Silin	68f2b5e7c7	add linker script to QA builds (#1030 )	2023-11-08 17:53:45 -08:00
arai713	3af8c81a72	Transpose 3d (#984 ) * added working example for 5D input using 1D kernel * example with 5D input tensor and 2d kernel - not working: issues with arguments * added updated version of 3d device op - changed descriptors/dims * added example file to check kernel * fixed descriptor and isSupportedArgument stride problem * added and modified kernel for 3d - updated tids/loop * adding some more 5d example files * fixed some issues * changes made for testing * working version: fixed error in stride for A, still a bit inefficient * cleaned up formatting/comments * updating formatting * more formatting fixes * fixing cmake, adding back gpu targets in cmake script * adding client example * added instances for client example * fixed errors in client example * implemented client ex with device_elementwise.hpp and device_elementwise_3d_impl.hpp * removed extra files * minor formatting and naming fixes * adding test files and profiler * fixing minor error * minor fix * removed unneccesary comments, renamed files * updated instance list for client example, added different layout example * removing instances * fixed error in instance generation * remove comments * update profiler and client example tensor layouts * fixed errors in test/profiler * updated vector dim access to enable vector load * updated test/profiler files * updated example with 1d kernel * updating profiler * renamed files --------- Co-authored-by: Jing Zhang <jizha@amd.com>	2023-11-08 19:45:07 -06:00
rocking	a3d9a2cd42	Layernorm4d (#1022 ) * Rename folder * Add layernorm 4d fwd example * Rename original layernorm example * Add layernorm 4d f16 test * Add layernorm4d_fwd client example * Support layernorm4D in ckProfiler * Rename groupnorm to groupnorm fwd in example * Rename layernorm and group fwd in test * Rename normalization to normalization_fwd (instances) * Add fwd to DeviceNormalization * Rename external api header * Rename folder, because we can also add bwd in this folder * Add fwd in layernorm and groupnorm (profiler * Fix compile error --------- Co-authored-by: Po Yen Chen <PoYen.Chen@amd.com>	2023-11-09 08:34:51 +08:00
Illia Silin	ce52621123	Support fp64 contraction on gfx94x. (#1029 ) * enable contraction fp64 on gfx94* * fix the logic	2023-11-08 15:03:18 -08:00
zjing14	98fd41f597	Add Gemm instances for performance improvement (#1018 ) * improve kpad * more tuning parameters * f16_f8_fp16 * cut test time * add f16_f8_fp16 * add f16_f8_f16 * testing instances for skinny cases * format * clean * add fp16_f8_fp16 * clang-format * add grouped gemm instalces * fixed profile grouped_gemm * clean * clean * clean * clean * clean * add missing instance func * fixed inferface --------- Co-authored-by: Jing Zhang <jizha@amd.com> Co-authored-by: root <root@sh5-1e707-rc06-38.mkm.dcgpu>	2023-11-07 09:09:58 -06:00
Daming Feng	aa0b979887	Add compute type check for convolution instances (#1015 ) * add compute type check for fp16 in forward convolution instances * Add compute type check for default compute types --------- Co-authored-by: Bartlomiej Kocot <barkocot@amd.com>	2023-11-07 01:33:11 +01:00
Illia Silin	b0568b728b	switch the hipTensor testing from mainline to develop branch (#1025 )	2023-11-03 12:36:41 -07:00
Bartlomiej Wroblewski	16eb824c90	Add missing ComputeDatatype in contraction_multi_ABD_xdl_fp16 (#1024 )	2023-11-03 08:22:11 -07:00
Bartlomiej Wroblewski	4ef704d8a6	Add support for mixed precision in contraction scale and bilinear (#973 ) * Add support for mixed precision in contraction scale and bilinear (#936) * Extract common functionality to separate files * Reference contraction: Remove incorrect consts from type_converts * Reference contraction: Add missing type_convert for dst value * Reference contraction: Fix incorrect order of B matrix dimensions * Add support for mixed precision in contraction scale and bilinear * Move using statements from instances to a common file * Move using statements from examples to a common file * Fix the order of B matrix dimensions across examples and profiler * Fix the computation of error threshold * Make ComputeDataType an optional argument * Include possible DataType -> ComputeDataType casting error in the threshold * Remove commented code * Make the ComputeDataType an optional argument in instance --------- Co-authored-by: Illia Silin <98187287+illsilin@users.noreply.github.com>	2023-11-02 14:26:33 -07:00
dependabot[bot]	73743aa0aa	Bump rocm-docs-core from 0.24.0 to 0.26.0 in /docs/sphinx (#987 ) Bumps [rocm-docs-core](https://github.com/RadeonOpenCompute/rocm-docs-core) from 0.24.0 to 0.26.0. - [Release notes](https://github.com/RadeonOpenCompute/rocm-docs-core/releases) - [Changelog](https://github.com/RadeonOpenCompute/rocm-docs-core/blob/develop/CHANGELOG.md) - [Commits](https://github.com/RadeonOpenCompute/rocm-docs-core/compare/v0.24.0...v0.26.0) --- updated-dependencies: - dependency-name: rocm-docs-core dependency-type: direct:production update-type: version-update:semver-minor ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>	2023-11-02 13:59:04 -07:00
Bartłomiej Kocot	f27ea94ecb	Add ScaleAddScaleAddRelu post op for conv fwd (#1006 ) * Add ScaleAddScaleAddRelu post op for conv fwd * Fixes * Fix instance file name * Minor fix	2023-11-01 18:31:30 -05:00

1 2 3 4 5 ...

1114 Commits