composable_kernel

mirror of https://github.com/ROCm/composable_kernel.git synced 2026-07-14 11:07:44 +00:00

Author	SHA1	Message	Date
Yi DING	790e7bbd4b	Fix fp8 convert & add option for basic example (#2129 ) [ROCm/composable_kernel commit: `8add2cf45d`]	2025-04-27 16:26:05 -07:00
Po Yen Chen	7b61b0e4fa	Avoid using store_tile_raw() for fp32 tensors (#2072 ) [ROCm/composable_kernel commit: `3d4d70d2fc`]	2025-04-26 23:07:41 -07:00
joyeamd	c27b313602	SWDEV-52596 for hdim=256, when use splitkv pipeline, two new pipelines need to be added (#2126 ) [ROCm/composable_kernel commit: `41541aff7a`]	2025-04-25 16:31:09 +08:00
rocking	cb729623b5	Only generate specific hdim (#2120 ) [ROCm/composable_kernel commit: `02ce6d39ea`]	2025-04-24 18:52:58 +08:00
lalala-sh	7748793e09	Moe gemm activation (#2026 ) * fix useless code and remove usless oob * clang format * fix coredump in e2e test * fix2 * fix clang format * fix output oob * impl int64 but result not correct * int64 index ok now * input output all ok * fix uint32 * revert v1 test * use uint32 * mork to support 13w tokens * moe sorting fix moebuf * fix merge * update moe api fix aiter build * fix buid * fuse silu * silu ok * acale ok * add silu * change code * gemm2 ok * gufusion compatible ok, fix warnings * gu fusion for m32 m64 ok * support bf16 cshuffle * i4 gemm2 ok * i4 gemm2 ok and i4 gemm1 build * 16x16 run ok * change flops; change cshuffle dtype * fuse gelu silu act in moe gemm1 * fp8 with act ready * int4 act ready * remove useless changes * remove useless code change * fix clang format * add the arch limit of int4 moe gemm * fuse moe activation * fix fp8 16x16 * fix no quant case * fix bugs * fix fp8 gufusion bug * remove useless comments * refine activation code & complete moe example * fix int8 bugs * merge tkw1 --------- Co-authored-by: coderfeli <coderfeli@163.com> Co-authored-by: feli <felix.li@amd.com> Co-authored-by: illsilin <Illia.Silin@amd.com> Co-authored-by: root <root@hjbog-srdc-51.amd.com> Co-authored-by: Illia Silin <98187287+illsilin@users.noreply.github.com> [ROCm/composable_kernel commit: `39ba03f25d`]	2025-04-23 10:35:34 +08:00
Gino Lu	d20a94b709	[CK-Tile] warp-gemm support for using V_MFMA_F32_16x16x32_BF16 (#2073 ) * draft v_mfma_f32_16x16x32_bf16 * fix error config and add debug code. * Solve the CShuffle Problem * draft v_mfma_f32_16x16x32_bf16 * fix error config and add debug code. * Solve the CShuffle Problem * fix error while testing new command * Finished the feature of new mfma 161632 * Addressed the comment --------- Co-authored-by: ThomasNing <thomas.ning@amd.com> [ROCm/composable_kernel commit: `504f563f78`]	2025-04-22 15:52:36 -07:00
lalala-sh	2d0b5aba13	enable do top k weights in moe stage1 gemm (#2094 ) * add switch for mul topk weights * fix bf16/f16 bugs * complete [ROCm/composable_kernel commit: `bcf5bb41be`]	2025-04-18 10:45:49 +08:00
Andriy Roshchenko	348760d56e	MX GEMM - Add MX BF8 example (#2071 ) * Add MX GEMM example for MX BF8 * Verified MX FP8 with 16x16x128 scale builtin * Verify MX BF8 GEMM with BF16 output [ROCm/composable_kernel commit: `da54464cce`]	2025-04-16 15:25:02 -06:00
BingYuan.Zhou	4ec293cb4b	[flatmm] implement basic fp16 flatmm (#2089 ) * [flatmm] implement basic fp16 flatmm * fix CI build fail --------- Co-authored-by: root <root@hjbog-srdc-50.amd.com> Co-authored-by: solin <bingzhou@amd.com> [ROCm/composable_kernel commit: `eaf1f0bf3b`]	2025-04-16 16:51:17 +08:00
felix	4c44fa374e	add preshuffle gemm fp16 (#2036 ) * add preshuffle gemm fp16 * clang format and test ok * Update gemm_multiply_multiply_xdl_fp16_bpreshuffle.cpp remove useless comments in example * Update gemm_multiply_multiply_xdl_fp16_bpreshuffle.cpp remove 2 --------- Co-authored-by: coderfeli <coderfeli@163.com> [ROCm/composable_kernel commit: `c5975529bb`]	2025-04-16 10:53:21 +08:00
joyeamd	19a9980cd5	fmha hdim256 vectorize improve (#2086 ) For hdim 256, will not have vectorized buffer load when seqlen % 256 != 0 and hdim % 256 = 0; this commit tries to solve this condition. [ROCm/composable_kernel commit: `94d47b1680`]	2025-04-16 09:21:04 +08:00
Andriy Roshchenko	d9c9f17c3d	MX GEMM - New GEMM pipeline for MX data types (#2059 ) * Allow selection of mfma_scale instructions * Read B tensor from LDS to VGPR in chunks of 16 in MFMA order * Add constexpr and synchronize return type for `get_exponent_value` * Pass scales by reference and add comments to `mfma_scale_f32_32x32x64` * Add support for microscaling instructions in `XdlopsGemm` * Fix `mfma_scale_f32_16x16x128f8f6f4` wrapper * Remove software implementation of MX GEMM * Make interface of `intrin_mfma_scale_f32_16x16x128f8f6f4<16, 16>` consistent with the other scale instruction * Update README * Updated CHANGELOG * Remove unused static methods [ROCm/composable_kernel commit: `7106976a72`]	2025-04-15 17:17:07 -06:00
Mingtao Gu	e8db9f0220	CK pk_i4_t test failures fix (SWDEV-518629) (#2075 ) * fix pk_i4_v3 tests failures in Unbuntu env. * fix pk_i4_t tests failure on Unbuntu issues. * some fixed. --------- Co-authored-by: mtgu0705 <mtgu@amd.com> [ROCm/composable_kernel commit: `56378f810f`]	2025-04-14 16:58:57 +08:00
jakpiase	d76ebf9795	[CK_TILE] Add 2:4 structured sparsity support for fp16 gemm (#1957 ) * add structured sparsity fp16 support for gemm * added reviewer suggestions * update changelog * update changelog * add reviewers suggestions * Minor fix * clang fix * fix doxygen [ROCm/composable_kernel commit: `6c61f4d237`]	2025-04-11 12:18:26 +02:00
slippedJim	959225947a	add fmha fwd splitkv receipt for aiter c++ api (#2068 ) * add s_randval for c++ api * Fix bug of bias in splitkv --------- Co-authored-by: rocking <ChunYu.Lai@amd.com> [ROCm/composable_kernel commit: `5f885d2b7a`]	2025-04-10 23:21:13 +08:00
Illia Silin	7546e4bafe	enable gfx115x support (#2065 ) [ROCm/composable_kernel commit: `3e6d21adeb`]	2025-04-09 10:06:42 -07:00
slippedJim	753a84d9d5	Add new receipt (#2055 ) [ROCm/composable_kernel commit: `5a22b61de5`]	2025-04-07 14:18:01 +08:00
Thomas Ning	03b4c5322d	Add the MI355 support for CK TILE GEMM (#2046 ) * Get the root cause of the ck tile gemm failing on mi355 * Fix the ck tile gemm on MI355 * delete the debug info [ROCm/composable_kernel commit: `50d1f8ff90`]	2025-04-03 11:48:54 -07:00
aledudek	b7359bcfac	Post-merge changes for fully async args copy in ck grouped gemm (#1991 ) * Post-merge changes for fully async args copy in ck grouped gemm * Post-merge documentation and naming changes * Build fix and updated changelog * Revised comments [ROCm/composable_kernel commit: `9329432f6c`]	2025-04-03 13:35:43 +02:00
Muhammed Emin Ozturk	532127f25d	f8/bf16 GEMM Stream-K (#1879 ) [ROCm/composable_kernel commit: `dd4c12b155`]	2025-03-31 20:30:17 -06:00
rocking	01ea8aa249	Reduce redundant space in bias tensor (#2024 ) Co-authored-by: Po Yen Chen <PoYen.Chen@amd.com> [ROCm/composable_kernel commit: `8a20b62e91`]	2025-03-28 21:58:06 +08:00
felix	20ffa0f474	hotfix fix sorting int64 (#2025 ) * fix sorting int64 * clang format * fix example issue * update WA issue # --------- Co-authored-by: coderfeli <coderfeli@163.com> Co-authored-by: carlushuang <carlus.huang@amd.com> [ROCm/composable_kernel commit: `a82f338fb9`]	2025-03-28 11:31:52 +08:00
felix	900acdc2db	ckmoe: change cmake; use smaller shape for i4 (#2027 ) * change cmake; use smaller shape for i4 * fix pki4 run * fix typo * fix runtime arch logic for moe_gemm2 example --------- Co-authored-by: coderfeli <coderfeli@163.com> Co-authored-by: illsilin <Illia.Silin@amd.com> [ROCm/composable_kernel commit: `36d50de50e`]	2025-03-27 09:04:31 -07:00
Illia Silin	73a5a3c463	Disable all pk_i4 tests for all targets except gfx942/950. (#2022 ) * only build gemm_fp8_pk_i4 examples for gfx942/950 * fix cmake logic * moved the architecture check to IsSupported function * Revert "moved the architecture check to IsSupported function" This reverts commit `056d2a08b3`. * disable all pk_i4 tests for targets other than gfx942/950 * fix cmake logic [ROCm/composable_kernel commit: `23a949706c`]	2025-03-26 15:15:57 -07:00
Illia Silin	27de1d431a	Make sure gemm_fp8_pk_i4 examples only build and run on gfx942/950. (#2010 ) * only build gemm_fp8_pk_i4 examples for gfx942/950 * fix cmake logic * moved the architecture check to IsSupported function * Revert "moved the architecture check to IsSupported function" This reverts commit `056d2a08b3`. [ROCm/composable_kernel commit: `99b2bbc1d6`]	2025-03-25 14:43:38 -07:00
Andriy Roshchenko	75ef4c83bf	MX GEMM examples with FP8, FP16, and E8M0 scales (#2016 ) * Add `scalar_type` specification for E8M0 exponent * Specialize `nnvb_data_t_selector` for E8M0 exponent * Remove partial specializations for `scalar_type` of `non_native_vector_base` template * Reword command line helper string * Create MX GEMM examples for different scales [ROCm/composable_kernel commit: `72d888821c`]	2025-03-25 15:33:03 -06:00
ruanjm	ce1d20c2c6	[CK_TILE] Improve RMS/Layer Normalization 2 Pass Pipeline Performance (#1861 ) * 50ms -> 28ms * Fix bug in non fuse_add_store cases * Fine tuned setting for 2 pass pipeline * adjust workload * remove unnecessary change * add layernorm * Adding output quant and unquant results at the same time. * fix test * fix format * tune for cases 128x640 and 128x1024 * bug ifx [ROCm/composable_kernel commit: `d49abdaa87`]	2025-03-25 20:09:45 +08:00
Andriy Roshchenko	bbdd7f6d57	Introduce MX GEMM for FP8 data type (#2000 ) [ROCm/composable_kernel commit: `6660dc6b8e`]	2025-03-24 15:41:07 -06:00
carlushuang	e1122c5c27	add mask support in hdim=192/128 (#1999 ) [ROCm/composable_kernel commit: `6c08c5c46d`]	2025-03-21 18:28:43 +08:00
BingYuan.Zhou	c245d569d5	fix ck_tile/basic_gemm build error (#1988 ) [ROCm/composable_kernel commit: `5a0d693b86`]	2025-03-20 22:01:14 -07:00
felix	bd00da1848	change cmake (#2006 ) Co-authored-by: coderfeli <coderfeli@163.com> [ROCm/composable_kernel commit: `902dbe89ad`]	2025-03-20 19:25:11 -07:00
carlushuang	23340c5dd5	[CK_TILE] return value with macro in ck_tile::kernel_launch API (#1982 ) * return value with macro and revert the return value * [CK-TILE] no-macro launch api solution (#1992) * no-macro solution * address -Wcomma --------- Co-authored-by: Max Podkorytov <4273004+tenpercent@users.noreply.github.com> [ROCm/composable_kernel commit: `e3c9886cdf`]	2025-03-20 11:00:29 -07:00
jakpiase	f1262b783a	[CK_TILE] Switch to universal gemm for batched and grouped gemms (#1919 ) * switch to universal gemm for batched and grouped gemms * added reviewer comments * fixed grouped gemm tests [ROCm/composable_kernel commit: `0e91d32c61`]	2025-03-20 11:17:04 +01:00
rocking	b0f323c4ec	Sync the kname with instance name (#1989 ) Co-authored-by: Po Yen Chen <PoYen.Chen@amd.com> [ROCm/composable_kernel commit: `b819c217e4`]	2025-03-20 00:06:45 +08:00
Illia Silin	77ed99efe7	Add a daily CI build on gfx908. (#1987 ) * add one daily ci build on gfx908 * add redis invocation tag for gfx908 * make ci build for gfx908 conditional * fix groovy logic * add option to run perf tests for gfx908 * disable a few tests on mi100 [ROCm/composable_kernel commit: `1342ecf7fb`]	2025-03-17 18:08:53 -07:00
aledudek	73d207bd4e	Async grouped gemm v3 (#1940 ) * Fully async grouped gemm * Remove commented code * Remvoe maybe_unused * host kernel args * Checkpoint segfault debugging... * Working part1 * Working part2 * Remvoe comments... * Use void ptr for gemm kernel host args * Fix device_grouped_gemm_multiple_d_dl build issue * Fix device_grouped_gemm_xdl build issue [ROCm/composable_kernel commit: `5095906975`]	2025-03-17 16:42:43 +01:00
valarLip	99d7424a14	hotfix fmoe build issue (#1976 ) [ROCm/composable_kernel commit: `52b1cd7780`]	2025-03-13 15:11:59 +08:00
carlushuang	f2dd57b76f	Reapply "[CK_TILE] support hdim=192/128 pair for deepseekv3 (#1961 )" … (#1971 ) * Reapply "[CK_TILE] support hdim=192/128 pair for deepseekv3 (#1961)" (#1969) This reverts commit `b92caa3d84`. * fix codegen problem * Update config.hpp --------- Co-authored-by: Illia Silin <98187287+illsilin@users.noreply.github.com> [ROCm/composable_kernel commit: `3e81279d26`]	2025-03-13 11:41:39 +08:00
Illia Silin	7f849a89e3	disable tests that take too long to build for gfx90a (#1975 ) [ROCm/composable_kernel commit: `d4a6d69643`]	2025-03-12 17:54:03 -07:00
Illia Silin	b92caa3d84	Revert "[CK_TILE] support hdim=192/128 pair for deepseekv3 (#1961 )" (#1969 ) This reverts commit `45fbd9210a`. [ROCm/composable_kernel commit: `8cbcd3e0d0`]	2025-03-11 10:40:18 -07:00
Haocong WANG	1ed0b74c43	[Block Scale GEMM] Optimized block scale gemm (#1950 ) * Added two kernel for M=32 problem * Comment the first one * Enable multiply_multiply for Scale_Block_M = 1 for deepseek * Modify the a_thread offset since the A data load is different from B. * edit fp8 ab scale for Scale_Block_M=1 * edit GemmSpec to MNKPadding * enable blockwise pipelie v1 and v2. v1 is work for small K. * add instance for gemm_ab_scale * fix cmakelist of ckProfiler * optimize blockscale gemm. todo: reduce vgpr usage * fix a correctness bug * sanity checked * revert ckprofiler cmake changes * clang format * revert unnecessary changes. * remove commented codes. * split weight preshuffle library targets * bring back enable-post-misched=0 * fix build issues for gemm_multiply_multiply_fp8 instances * fix clang format * add verbose build flag when building for all targets * reduce path names for new instances * fix paths in cmake * refactor gemm_multiply_multiply library target * fix a bug in example * fix example 65 cmake * reduce the number of threads when building libs for all targets to 50 * use ninja to build for all targets * reduce teh number of threads when building for all targets * reduce the number of threads to 32 when building libs for all targets to 50 --------- Co-authored-by: mtgu0705 <mtgu@amd.com> Co-authored-by: chenjun <junchen2@amd.com> Co-authored-by: illsilin <Illia.Silin@amd.com> Co-authored-by: Illia Silin <98187287+illsilin@users.noreply.github.com> [ROCm/composable_kernel commit: `cbd74c2d12`]	2025-03-11 10:11:21 -07:00
Illia Silin	8fe2095d43	disable example_moe_gemm2_xdl_pk_i4 on gfx950 (#1968 ) [ROCm/composable_kernel commit: `aa42c3db06`]	2025-03-11 08:34:47 -07:00
carlushuang	45fbd9210a	[CK_TILE] support hdim=192/128 pair for deepseekv3 (#1961 ) * support hdim=192/128 pair * remove useless print * update [ROCm/composable_kernel commit: `7a93b16ff6`]	2025-03-11 21:07:40 +08:00
Mingtao Gu	fc98615212	Ck int4 moe develop (#1949 ) * Add Gemm fp8xint4 example and kernel, function pass. * Init Gemm_fp8xint4 Bpreshuffle * Added gemm_fp8xint4_Bpreshuffle files, function not checked yet * General fix. * fp8xint4 bpreshuffle function pass * fix. * init b preshuffle dequant in VGPR. * fix bug, function pass. * move b thread dequant copy to blockwise. * fix bug, function now passes. * modified the tile size to 256, 128x128x128. * fixed a bug. * Initial int4 moe, compile pass, function not check. * fix bug in moe_gemm1.cpp, now function pass. * test expert = 8 and function pass. * Added moe_pk_i4_gemm2, function pass. * Added b preshuffle pipeline v3 support. * fixed merge issue. fp8xint4 and fp8xint4_bpreshuffle function pass. * Split the blockwise pipeline for fp8xint4. * commit missing files * opt gemm2 to 2x2 wave * fix swizzle = false * update int4 moe with latest input changes. * update tile size. * enable pipeline v3. * fix nswizzle = true * commit a version for compiler debug. * Updated transfer_v3r1_gather to support pk_i4_t type. * for int4 moe2 for type_convert support. * remove some values between mfma instructions. * fix int4 moe * Updated transfer_v3r1_gather to support pk_i4_t type. * i4 support lds multiple shuffle * fixed int4 moe tflops calculation. * Modified CshuffleCShuffleMXdlPerWavePerShuffle to 1 to suit C multiple shuffle * updated gemm2. * change int4 moe example names * fix and format code. * format. * format codes. * update fp8xint4 example tile size. * add <unordered_map> header * fixed. * format. * Added conditional compilation for int4 -> fp8 conversion kernels --------- Co-authored-by: mtgu0705 <mtgu@amd.com> Co-authored-by: coderfeli <coderfeli@163.com> [ROCm/composable_kernel commit: `0db7c8f0b2`]	2025-03-10 11:16:44 +08:00
Thomas Ning	89f3ca4c89	Add the instance of MBlock=144 for GemmMultiplyMultiply (#1955 ) * tempsave, not selected * finish the feature and merge with develop --------- Co-authored-by: aska-0096 <haocwang@amd.com> [ROCm/composable_kernel commit: `c954bd0cfa`]	2025-03-07 13:44:06 -08:00
Max Podkorytov	9b160b318f	refactor ck-tile kernel launch (#1925 ) [ROCm/composable_kernel commit: `9e132eb77c`]	2025-03-07 08:29:40 -08:00
kylasa	676d236a5e	Addressing (Post Merge) code review comments for PR 1845 (#1883 ) * Addressing code review comments. * Addressing code review comments. * Reorganized code for better readability. * add ck_tile gemms for new types in CI * fix jenkins syntax * fix script syntax * Add the test cases back * Address the review comments * Address review comments * clang format * Solve the merging issues * Addressed the comments * clang format --------- Co-authored-by: illsilin <Illia.Silin@amd.com> Co-authored-by: ThomasNing <thomas.ning@amd.com> Co-authored-by: Adam Osewski <19374865+aosewski@users.noreply.github.com> [ROCm/composable_kernel commit: `66c5f5b0b6`]	2025-03-06 11:40:30 -08:00
Illia Silin	53b89436b2	remove support for gfx940 and gfx941 targets (#1944 ) * remove support for gfx940 and gfx941 targets * update changelog [ROCm/composable_kernel commit: `9b51c08bf7`]	2025-03-05 11:07:33 -08:00
feli	6fd94cff45	ck moe gemm implement (#1936 ) * port all moe changes from ck_moe_gemm branch * refine codes in the pr * fix tail odd * fix clang format * fix clang format2 * make hot loop scheduler compatible with 16x16 and 32x32 * clang format * fix per token quant * rename moe example * clang format --------- Co-authored-by: coderfeli <coderfeli@163.com> [ROCm/composable_kernel commit: `3786e16375`]	2025-03-05 15:56:55 +08:00
jefyang1	dfd15c220d	Remove CK_USE_AMD_MFMA_GFX950 (#1935 ) * Add runtime check in example_gemm_xdl_streamk for gfx950 * Add runtime check in grouped conv fwd examples for gfx950 * Disable CK_USE_AMD_MFMA_GFX950 * Add new instances for gfx950 * Fix test_gemm_universal on gfx950 [ROCm/composable_kernel commit: `c95bda93ba`]	2025-03-04 10:32:25 -08:00

1 2 3 4 5 ...

526 Commits