composable_kernel

mirror of https://github.com/ROCm/composable_kernel.git synced 2026-05-16 19:09:59 +00:00

Author	SHA1	Message	Date
Illia Silin	3609ff10f7	Refactoring cmake files to build data types separately. (#932 ) * refactor cmake files for the tests * refactor cmake files for examples * fix cmake for gemm example * fix the cmake file for all examples * add splitting by data types in gemm_splitk instance header * rename test to reflect only dl instances are used * clean up CI workspace, update cmake for instances * change the jenkinsfile syntax * build all instances except DL on gfx11 * move workspace cleanup after stages * clean up workspace after every stage * isolate data types in grouped_conv_fwd header * isolate dl instances for grouped_conv2d_fwd * fix syntax * fix cmake and batchnorm instances * fix typo * fix reduction instances * fix grouped_conv headers * fix syntax * replace parsing logic for instances, replace bfp16 with bf16 * fix the client examples build * clean up DTYPES from instances cmake files * update the parsing logic in cmake files * make an exception for reduction kernels * update few remaining cmake files to handle DTYPES * fix syntax * fix cmake conflicts * replace f8 with fp8 test name * resolve conflicts for dpp instances [ROCm/composable_kernel commit: `bba085d2b5`]	2023-09-20 22:15:56 -07:00
Bartlomiej Wroblewski	4497a8874f	Fix DL GEMM instances with too large vector size (#901 ) * Fix vector lengths of DL GEMM instances with padding * Add checks for correctness of vector lenghts in DL GEMM [ROCm/composable_kernel commit: `63cd459248`]	2023-09-18 14:08:23 +02:00
Bartlomiej Wroblewski	b4064d1401	Add new instances and support for small cases in DPP8 GEMM (#896 ) [ROCm/composable_kernel commit: `547dbcfbc2`]	2023-09-12 10:05:23 -05:00
Bartlomiej Wroblewski	02f8f707e8	Redesign the DPP8 GEMM kernel to use warp-wise component (#863 ) * Redesign the DPP8 GEMM kernel to use warp-wise component * Review: Improve error messages * Review: Remove unnecessary empty lines * Review: Fix M, N per thread names * Review: Rename mfma_input_type to dpp_input_type * Review: Fix tensor adaptor; remove unnecessary element * Review: Remove calls to dpp_gemm's MakeCDescriptor * Review: Add blockwise doc, change function names to include dimension names * Review: Remove duplicated code; Move Block2CtileMap alias to the top of the file * Review: Add __restrict__ keywords * Review: Use MatrixPadder for padding A, B, C matrices * Review: Remove hardcoded datatypes * Review: Change names from FloatX to XDataType * Review: Introduce AK0 and BK0 instead of a single K0 * Review: Remove construction of dpp_datatypes object * Review: Rename DppInstrRunner to DppLanegroupGemm [ROCm/composable_kernel commit: `37a8c1f756`]	2023-09-06 11:44:09 -05:00
Jun Liu	2fb9a37881	[HotFix] add config and version files to pass on build info (#856 ) * experiment with config file * experiment with version.h config * add more info to version.h * minor updates * minor updates * fix case where DTYPE is not used * large amount of files but minor changes * remove white space * minor changes to add more MACROs * fix cmakedefine01 * fix issue with CK internal conflict * fix define and define value * fix clang-format * fix formatting issue * experiment with cmake * clang format v12 to be consistent with miopen * avoid clang-format for config file [ROCm/composable_kernel commit: `c8a8385fdd`]	2023-08-23 11:36:17 -07:00
Bartlomiej Wroblewski	d4888118a5	Implement DPP8 based GEMM for Navi21 (#826 ) [ROCm/composable_kernel commit: `d4c84256f7`]	2023-08-14 15:46:27 -05:00
Po Yen Chen	b6e54f589e	Update tuning parameter & compilation options of DeviceGemmXdl<> instance (layout=TT) (#819 ) * Enable pipeline v2 opt for layout=TT instance * Use better thread mapping for reading A tile * Conditionally enable pipeline v2 opt * Allow enabling only fp16 gemm instances in profiler * Fix formatting error * Fix compilation error if we enable fp32 in profiler [ROCm/composable_kernel commit: `f7cc8c3b03`]	2023-08-02 10:32:22 -05:00
Illia Silin	f83a1c84c3	Disable DL kernels by default. (#816 ) [ROCm/composable_kernel commit: `9195435c77`]	2023-07-26 11:06:45 -05:00
Po Yen Chen	b49076fed0	Speed-up global memory reading for GEMM instances (#813 ) * Use better ThreadClusterLengths to speed up * Update B tile reading pattern for layout=NN instance [ROCm/composable_kernel commit: `f4ea560112`]	2023-07-25 18:54:47 -05:00
Illia Silin	74c83ffe26	Add mechanism to build CK for select data types, add Navi3x CI. (#790 ) * allow building CK for specific data types * add CI build and test stage on Naiv3x without some int8 instances * add missing gemm fp16 instances * add the changes to the missed cmake file * add empty lines at end of source files * Do not build quantization client example on navi3 in CI * disable batched_gemm_multi_d_int8 instances with DTYPES * disable device_conv2d_bwd_data_instance with DTYPES * fix ckprofiler for conv_bwd_data for int8 * properly isolate the conv_bwd_data int8 instances * remove empty line [ROCm/composable_kernel commit: `189ea3b9aa`]	2023-07-17 18:02:42 -07:00
Po Yen Chen	aff6040b5b	Split GEMM instance library & enable pipeline v2 optimization (#783 ) * Move source file into sub-directories * Add missing include directive * Split DeviceGemmXdl<> fp16 instances * Fix format * Remove unnecessary CMakeLists.txt * Add macros to toggle new features * Remove debug message * Turn off GEMM v2 pipeline optimization by default * Fix format * Extract duplicated string as list * Enlarge indent in CMakeLists.txt [ROCm/composable_kernel commit: `850144a0d3`]	2023-07-06 10:59:35 -05:00
Illia Silin	d40b8d5e2c	update copyright headers (#726 ) [ROCm/composable_kernel commit: `b94fd0b227`]	2023-05-31 18:46:57 -05:00
Bartłomiej Kocot	18002ddb3c	Add instances for fp16/int8 Gemm kernels (Navi21) (#717 ) * Add instances for fp16/int8 Gemm kernels (Navi21) * Extend instances with smaller tiles * Fix SrcVectorTensor for km_kn_mn int8 [ROCm/composable_kernel commit: `c2d7a29dec`]	2023-05-30 07:07:17 -05:00
Rostyslav Geyyer	ebf3f8571d	Add padding device_gemm_xdl instances (#529 ) Co-authored-by: Rosty Geyyer <rosty.geyyer@amd.com> Co-authored-by: Chao Liu <chao.liu2@amd.com> [ROCm/composable_kernel commit: `c7a4d36147`]	2022-12-07 17:46:03 -06:00
Rostyslav Geyyer	deb5b07204	Add pipeline v1/v2 selector, add more instances (#381 ) * Add gridwise gemm pipeline v1/v2 selector * Pipeline selector working, test-wise add pipeline options to one instance * Add gemm instances * Add debug info to DeviceGemmXdl * Add debug info to DeviceGemmXdl_CShuffle * Add debug info to DeviceGemmXdl_CShuffle and instances to gemm_add_add_fastgelu * Minor fix * Add debug info to DeviceBatchedGemmXdl and instances to batched_gemm * set up inter-wave configuration * use defualt loop scheduling for supported gemm ops for blanket-applying interwave scheduling for all supported gemm ops, define macro CK_EXPERIMENTAL_DEFAULT_TO_INTER_WAVE_SCHEDULING=1. this should be discouraged though as it is not covered by CI * Add enum PipelineVersion * Update instances * Format * Fix the merge conflict * Add flags to disable added instances * Test disable flag check * Disable flag check * Enable the instances Co-authored-by: Anthony Chang <ac.chang@outlook.com> [ROCm/composable_kernel commit: `1a0b0e7bec`]	2022-11-02 16:50:48 -06:00
Adam Osewski	8a8f8521f9	Refactor device op implementations into `impl` subdirectory. (#420 ) * Move kernel implementation files under impl directory. * Update examples paths. * Update device kernel impl include paths. * Update tensor operation instances include paths. * Update profiler and tests include paths. * Clang-format * Update include paths for batched gemm reduce * Refactor UnitTest ConvNDBwdWeight. * Refactor fwd and bwd data convND UT. * Fix used test macro. * Fix include path. * Fix include paths. * Fix include paths in profiler and tests. * Fix include paths. Co-authored-by: Adam Osewski <aosewski@amd.com> [ROCm/composable_kernel commit: `3048028897`]	2022-10-13 09:05:08 -05:00
cloudhan	91f93c0e19	Change all device operations to use add_instance_library (#338 ) * Change all device operations to use add_instance_library to avoid duplicated cmake configuration. * update DeviceMem Co-authored-by: Chao Liu <chao.liu2@amd.com> [ROCm/composable_kernel commit: `fb1cbf025b`]	2022-08-13 12:17:58 -05:00
Chao Liu	82745bffde	N-D Tensor Contraction example, instance, and client example (#270 ) * adding contraction * add contraction example * update examle * update example * format * update readme * clean header * clean header * contraction with multiple D * rename * fix naming issue; add instances for contraction+bilinear * change assumed virtual layout of contraction; add client example * update example * update * contraction+scale * use type_convert * rename [ROCm/composable_kernel commit: `4fe9c393b8`]	2022-07-07 14:31:11 -05:00
Chao Liu	4be57e5afa	Gemm+Bilinear (#316 ) * refactor * update example * update example * gemm bilinear * clean * update [ROCm/composable_kernel commit: `9e4429f9c3`]	2022-07-02 09:15:38 -05:00
Chao Liu	74b6e85eaf	Improve external interface for GEMM and GEMM+add+add+fastgelu (#311 ) * interface for GEMM and GEMM+add+add+fastgelu * rename namespace * instance factory * fix build * fix build; add GEMM client example * clean [ROCm/composable_kernel commit: `0dcb3496cf`]	2022-06-30 22:11:00 -05:00
Chao Liu	675e7b7956	External Interface (#304 ) * add client example * clean * clean * reorg * clean up profiler * reorg * clea * fix profiler * function for getinstances * update client example * update client example * update client example * update * update example * update Jenkins file * update cmake * update Jenkins [ROCm/composable_kernel commit: `aebd211c36`]	2022-06-26 19:39:02 -05:00
Chao Liu	2ef299e0ad	add license in file (#303 ) [ROCm/composable_kernel commit: `d3051d7517`]	2022-06-24 23:32:43 -05:00
Chao Liu	9df0a11a51	Absolute include path (#281 ) * ad gelu and fast_gelu * added GeLU and fast GeLU * clean up * add gemm+fastgelu example * add gemm+gelu instances * update profiler * clean up * clean up * adding gemm+bias+activation * clean * adding bias * clean * adding gemm multiple d * debugging * add gemm bias add fastgelu * rename, clean * refactoring; add readme * refactor * refactor * refactor * refactor * refactor * refactor * fix * fix * update example * update example * rename * update example * add ckProfiler * clean * clean * clean * clean * add client app example * update readme * delete obselete files * remove old client app * delete old file * cleaning * clean * remove half * fix header path * fix header path * fix header path * fix header path * fix header path * fix header path for all examples * fix header path * fix header path * fix header path * fix header path * fix header path * fix header path * fix header path * fix header path * fix header path * revert client app example * clean build * fix build * temporary disable client test on Jenkins * clean * clean * clean [ROCm/composable_kernel commit: `d1db6a0c3e`]	2022-06-24 20:51:04 -05:00
ltqin	d1f7ed99ec	Add FP64 XDL GEMM built-in function (#199 ) * add intrin_mfma_f64_16x16x4f64 * add example * gemm reference add double data type * chang init data * fix M N PerXdlops * fix ifdef * add comparsion config * add conv fwd example * format log out * change rc matrix egister layout * reorganize example * reorganize example 2 * format,because merge develop * fix call impl adding acc data type * lost ; * add compiler warning * change example tunning parameters * add test for fp64 * add instance * add test/gemm/gemm_fp64.cpp * fix get name issue * remove some tunning parameter * fix conflict * format * use integer value for GEMM test * add acc data type * remove typeid because fp16 * fix streamconfig etc bug from merging develop * format * remove test_gemm_xdl_fp64 * add AccDataType * AccDataType problem Co-authored-by: qinletao <letaoqin@amd.com> Co-authored-by: Chao Liu <chao.liu2@amd.com> [ROCm/composable_kernel commit: `3e6c2610ae`]	2022-05-26 14:48:57 -05:00
Jianfeng Yan	050fc62872	Navi21 gemm (#197 ) * start adding navi21 GEMM * navi_gemm_km_kn_mn_fp32 compiles and passes one test. * rename variables and functions in gridwise_gemm_dlops_v1r3 * add other 3 layouts; format instance * adding more tuning parameters add tuning parameters for other 3 layouts * add gemm_dlops_f16 * tmp * add dependence of DeviceGemm::IsSupportedArg() on arch * minor changes * minor changes * minor changes * minor changes * minor changes * minor changes * minor changes * push gemm_dlops into profiler * minor changes * if using xdl or dlops is moved into profiler_gemm_impl * minor changes * minor changes * remove is_xdl from profile_gemm_impl * make IsSupportedArg dependent on arch for other device_gemm * minor changes * minor changes * fix a bug in f_generate_tensor_value * add 64x64x64 for gemm_dlops_int8 * add 64x64x64 for gemm_dlops_int8 * comment out 3 layouts in gemm_dlops_int8; add 32x32x32 for gemm_dlops_int8; init A values to 1 * fix * start fixing tuning parameters * monir * minor changes * minor changes * minor changes * fixing * adding example * adding example * adding example * add gemm fp32 example * clean up * use 128x128x16 as MNK tile in navi21 gemm example * bug fix * fix test * use new block c tile * clean * fix build Co-authored-by: Chao Liu <chao.liu2@amd.com> Co-authored-by: shaojiewang <wsjmessi@163.com> [ROCm/composable_kernel commit: `40b59a63cc`]	2022-05-24 12:19:27 -05:00
JD	569dd9f47b	Add host API (#220 ) * Add host API * manually rebase on develop * clean * manually rebase on develop * exclude tests from all target * address review comments * update client app name * fix missing lib name * clang-format update * refactor * refactor * refactor * refactor * refactor * fix test issue * refactor * refactor * refactor * upate cmake and readme Co-authored-by: Chao Liu <chao.liu2@amd.com> [ROCm/composable_kernel commit: `cec69bc3bc`]	2022-05-12 09:21:01 -05:00
Chao Liu	a5ad59ed11	Code refactor (#175 ) * format * improving pipeline * fix typo * format * adding thread group * adding thread group * adding thread group * adding gemm pipeline * tweak * refactor * refactor * add missing type convert * refactor * refactor * refactor * clean * fix build * refactor * format * clean up * use remove_cvref_t * clean * clean up * clean up * clean up [ROCm/composable_kernel commit: `ec7c2e912e`]	2022-05-09 14:57:59 -05:00
Anthony Chang	f9a5880af6	profiler: fix fp32 c-shuffle gemm tuning parameter (#194 ) [ROCm/composable_kernel commit: `7c0b149811`]	2022-04-22 15:48:51 -05:00
Chao Liu	5aa380eb6f	fix build (#171 ) [ROCm/composable_kernel commit: `646878162b`]	2022-03-31 20:30:20 -05:00
Anthony Chang	1450193e62	Tune & add conflict-free LDS gemm kernels (#159 ) * retune & add conflict-free bf16/fp16 c-shuffle gemm instances amend wrong K1 value in some fp16/bf16 kernel instances * make gemm cshuffle's timing behavior consistent with all other functions * clang-format * retune & add conflict-free fp32 c-shuffle gemm instances * retune & add conflict-free int8 c-shuffle gemm instances * update the underlying gridwise gemm of all c-shuffle gemm kernels * typo [ROCm/composable_kernel commit: `7db48f9008`]	2022-03-31 12:58:41 -05:00
Chao Liu	3f732cceab	Compile for gfx908 and gfx90a (#130 ) * adding compilation for multiple targets * fix build * clean * update Jekinsfile * update readme * update Jenkins * use ck::half_t instead of ushort for bf16 * rename enum classes * clean * rename * clean [ROCm/composable_kernel commit: `cd167e492a`]	2022-03-31 12:33:34 -05:00
Chao Liu	040a21aa38	clean (#143 ) [ROCm/composable_kernel commit: `2206136628`]	2022-03-22 21:55:03 -05:00
rocking5566	066110a454	Gemm_c_shuffle (4 layouts) X (fp32 bf16 int8) (#131 ) * [What] Separate fixpoint gemm from gemm example [Why] let example of gemm_int8 be pure gemm. [What] 1. Add gemm_requant_relu_requant, 2. Let CDataType be int32 in pure gemm, because no one use int8 CDataType. It is also part of gemm_requant_relu_requant * Fix path * Revise cmakelist due to merge develop * Add gemm fp16 test * Extract PrepareGemmTensor * Extract TestGemm * Add test for different layout * Add 4 layouts of shuffle version of fp32 * Add 4 layouts of shuffle version of int8 * Add 4 layouts of shuffle version of bf16 * replace all DeviceGemmPtr_ with DeviceGemmNoOpPtr to fit naming convension * Add test for non-shuffle verstion of gemm * Fix typo * Print kernel information * Add rest of the fp32 kernel to the test * 1. Add rest of the fp16 device iop. 2. Mark the invalid device operation Co-authored-by: rocking <chunylai@amd.com> [ROCm/composable_kernel commit: `485ea46a40`]	2022-03-21 15:59:51 -05:00
Chao Liu	82ad74304e	Reorganize files, Part 1 (#119 ) * delete obselete files * move files * build * update cmake * update cmake * fix build * reorg examples * update cmake for example and test [ROCm/composable_kernel commit: `5d37d7bff4`]	2022-03-08 21:46:36 -06:00

34 Commits