composable_kernel

mirror of https://github.com/ROCm/composable_kernel.git synced 2026-06-07 00:04:37 +00:00

Author	SHA1	Message	Date
Kevin Choi	75c4b9372f	remove noinline attr as it causes a lot more s_waitcnt's	2025-08-16 08:15:23 +00:00
Kevin Choi	598e3fec41	remove innerloop, move restrict parameters to mainloop and add noinline attribute.	2025-08-14 12:11:17 +00:00
Kevin Choi	3340408537	Create inner lambda with restrict parameters, add restrict to some parameters	2025-08-14 07:06:51 +00:00
aska-0096	3bc45ecbc7	save for debug	2025-08-14 03:43:54 +00:00
aska-0096	108abf00e0	Merge branch 'develop' of https://github.com/ROCm/composable_kernel into wip-async-tr-fa	2025-08-13 02:14:26 +00:00
Thrupti Raj Lakshmana Gowda	3f57ec3d2d	GEMM Multi D for CK Tile Engine (#2660 ) * Readme for GEMM Multi D * GEMM Multi D partial Progress * GEMM Multi D partial Progress! * CK Tile Engine GEMM Multi D : All Python files generated * Partial Progress * Partial Progress * Partial Progress * Partial Progress : Incorrect Result * Partial Progress : Debugging * Partial Progress : Correct Results * Partial Progress - Incorrect Results * Partial Progress - Commenting Passthrough bypass logic * Changing Passthrough to MultiplyMultiply * Correct Results! * Fix and debug the pass through feature * Sample commit * Correct Results : MultiplyMultiply * Code Cleanup * Removing Failed Instances * Working code before Unary element support * Custom Elementwise Function support and working implementation for Mul and Add * Updating README * Working for Passthrough * Review Comments : Minor Fixes * Review Comments : Minor Fixes * Readme Updated * Partial Changes after Rebase * Working Code : Changes after Rebase * Updating Jenkins file * Removing default value changed while testing * Configuration changes in config files * Tile Handler changes in GEMM Multi D Tile Engine * Tile Handler changes in GEMM Multi D Example * Change log for Gemm Multi D in CK Tile Engine * Configuration changes in config files --------- Co-authored-by: ThomasNing <thomasning@amd.com>	2025-08-12 16:05:05 -07:00
joyeamd	0856b3f4a2	[CK_TILE]fix ck_tile's moe_sorting example in gfx11 (#2667 ) * fix ck_tile's moe_sorting example in gfx11 * fix clang format --------- Co-authored-by: illsilin_amdeng <Illia.Silin@amd.com>	2025-08-12 12:33:56 -07:00
aska-0096	0810799e25	refactor blockgemm change, isolate to v2;	2025-08-12 14:25:50 +00:00
asleepzzz	5b39de4bb6	Revert "Optimize fmha fwd decode & prefill for gfx950 (#2641 )" (#2670 ) This reverts commit `b7322a521a`.	2025-08-12 20:27:10 +08:00
Haocong WANG	b7322a521a	Optimize fmha fwd decode & prefill for gfx950 (#2641 ) * Fix for fwd/bwd kernel build filter * fix bwd code * save an example for __bf16 type * temp save, waiting for debug * tempsave, fmha_decode * temp save, change all instance to 1wave * fix async copytest bug * Add block_sync_lds_direct_load utility * fix the s_waitcnt_imm calculation * Improve s_waitcnt_imm calculation * fix vmcnt shift * add input validation and bug fix * remove unnecessary output * move test_copy into test * temp save * tempsave * compile pass * tempsave, trload+asyncload done * tempsave. asynccopy+trload sanity checked * remove unnecessary features * fix the lds alignment caused performance regression * enable prefill overload operator(). * remove all lds bankconflict with xor layouts * enable larger tile size; upgrade xor pattern * upgrade prefill pipeline; simple iglp; consistent data produce and consume order * small refactor * Load Q through lds, implement xor; * add vmcnt guard before load ktile * Add v_permlaneb32 for block_reduce. Disable it as it will cause un-coexecutable packed math in FA * Add XOR fold strategy for hdim<128, but perf dropped; disable it by default; wait further perf debug * add __restrict__ to tr load * merge fa_decode pipeline into fmha_fwd api * remove unnecessary files; rename some files * Remove unnecessary changes * bug fix, clang format; * remove non-necessary change * fix clangformat with 18.1.3 * fix bugs * fix bug * fix bug on non-gfx950 * fix bugs in gemm * fix bug in pki4 * tempsave, update the blocksync functions * change the warp setting for hdim32 fmha fwd * clang format * fix conflict. disable all v-col instance for fmha fwd * Fix the bug * clang format --------- Co-authored-by: Max Podkorytov <4273004+tenpercent@users.noreply.github.com>	2025-08-12 19:43:14 +08:00
aska-0096	75f6f6bac4	Merge branch 'develop' of https://github.com/ROCm/composable_kernel into wip-async-tr-fa	2025-08-12 09:04:41 +00:00
Yi DING	8e1eb0c1ee	[CK_TILE] FMHA BWD Decode Pipeline (#2643 ) * Fix distr * Duplicate block_fmha_bwd_dq_dk_dv_pipeline_trload_kr_ktr_vr * decode 16x16 o2	2025-08-12 17:02:52 +08:00
aska-0096	bcc05eee62	Fix the bug	2025-08-12 08:46:06 +00:00
aska-0096	96d24497f5	fix conflict. disable all v-col instance for fmha fwd	2025-08-12 04:02:41 +00:00
aska-0096	1716171be4	Merge branch 'develop' of https://github.com/ROCm/composable_kernel into wip-async-tr-fa	2025-08-12 03:52:34 +00:00
Yi DING	4fde1646e5	[CK_TILE] FMHA BWD Optimization For GFX950 (#2628 ) * simplify fmha_bwd_kernel MakeKargs & dq_dram_window * simply duplicate * trload pipeline * Try two-stage * add prefetch * optimize & iglp	2025-08-12 11:11:55 +08:00
aska-0096	1c98007901	clang format	2025-08-12 01:53:31 +00:00
aska-0096	3868ddd708	Merge branch 'develop' of https://github.com/ROCm/composable_kernel into wip-async-tr-fa	2025-08-11 15:59:40 +00:00
aska-0096	b86f7786e2	tempsave, update the blocksync functions	2025-08-11 14:21:09 +00:00
Yashvardhan Agarwal	191c62967b	Fixes to "General 2D Reduction Kernel" (#2535 ) (#2656 ) * fix reduce2d - revret the combine_partial_results() chnages - remove auto from function def * clang-format	2025-08-11 15:01:33 +02:00
aska-0096	7b8052d7ca	fix bug in pki4	2025-08-10 06:00:51 +00:00
aska-0096	76cbbb84a2	fix bugs in gemm	2025-08-09 03:25:12 +00:00
aska-0096	efb8549279	fix bug	2025-08-08 17:53:19 +00:00
aska-0096	729e8785fb	fix bugs	2025-08-08 15:42:15 +00:00
aska-0096	250dc13c75	fix clangformat with 18.1.3	2025-08-08 09:31:01 +00:00
aska-0096	78edd7303b	bug fix, clang format;	2025-08-08 09:04:02 +00:00
aska-0096	3b9fb6af38	Remove unnecessary changes	2025-08-08 08:08:03 +00:00
aska-0096	6bb57c2c57	Merge branch 'develop' of https://github.com/ROCm/composable_kernel into wip-async-tr-fa	2025-08-08 07:50:12 +00:00
aska-0096	1ecee378d5	remove unnecessary files; rename some files	2025-08-08 06:19:31 +00:00
aska-0096	b4640a9de6	merge fa_decode pipeline into fmha_fwd api	2025-08-08 05:46:18 +00:00
Max Podkorytov	ab26026835	[CK-tile] add more tests for batched transpose testing the rectangular block tile sizes (#2634 ) * add failing tests * swap out and reference * add constraint assert to transpose input distribution * test both pipelines with rectangular block tile * print mismatched indices * add a smaller failing test for old pipeline * print grid and block * fill output before operating on it * swap m/n tile sizes and make one test pass * add device syncs * add one more flipped test case * flip block tile at host arg init * fix tiles for lds pipeline * clang-format * rename tests * roll back error check * remove device syncs * reduce large test case's size	2025-08-07 16:51:53 -07:00
Gino Lu	5d6d236b25	Add e8m0 scaled convert into CK_TILE (#2617 ) * first commit * remove redundent code * modify according to comments. * fix type_convert error with scaled_type_convert	2025-08-07 21:37:28 +08:00
Yi DING	b0a97498b0	[CK_TILE] FMHA BWD Remove Unnecessary Padding (#2550 ) * Remove unnecessary pssk * Add BlockFmhaBwdDQDKDVPipeline wrapper * Resolve copilot comments & Remove kpad & fix * Remove spad	2025-08-07 21:24:43 +08:00
Sami Remes	ffdee5e774	[CK_TILE] Enable printing more structures in CK-Tile (#2443 ) * Add more printing to core cktile * Revert other changes in static encoding pattern * Refactor to using a free print() function * Remove loops and print just the containers * Print tuple with better formatting, fix sequence compilation * Add some tests for print utility * Add print utility header * Print for static_encoding_pattern * add buffer_view printing * Align vector_traits * Fix formatting * Lower-case enum strings Co-authored-by: Christopher Millette <63608002+cgmillette@users.noreply.github.com> * Remove empty comment lines * Fix test with lower-case too * Reduce repeated code in print tests, move helper function closer to type definition, test X&Y * Add test_print_common.hpp * add print.hpp in core.hpp --------- Co-authored-by: Aviral Goel <aviral.goel@amd.com> Co-authored-by: Christopher Millette <63608002+cgmillette@users.noreply.github.com> Co-authored-by: Adam Osewski <19374865+aosewski@users.noreply.github.com>	2025-08-07 15:45:27 +03:00
Enrico Degregori	21e9983913	Revert "Add padding to 1x1Stride1Pad0 conv specialization (grouped conv bwd weight) (#2610 )" (#2637 ) This reverts commit `2203b0ddfe`. Co-authored-by: Bartłomiej Kocot <barkocot@amd.com>	2025-08-07 12:30:08 +02:00
Bartłomiej Kocot	5328b232b2	Grouped Convolution Forward Infer Bias Bnorm Activ (#2621 ) * Grouped Convolution Forward Infer Bias Bnorm Activ * 3d	2025-08-07 08:36:47 +02:00
Yashvardhan Agarwal	4750b293fe	General 2D Reduction Kernel (#2535 ) * General 2D Reduction Kernel * Move the reduction kernel from the example * Split the code and add the necessary policy, problem, shape files as per ck_tile convention * Add/modify the headers * Modified the example to work with the 'new' kernel * Added tests for the kernel * N-D refernce reduce * Added support for N-D input with transform to 2D * Added padding to support various input sized tensors * Bug fix in the thread buffer constructor * Some comments to explain the reduce2d block kernel * comments resolution * clang-format * comments resolution * clang-format * clang-format * comments resolution * clang-format	2025-08-06 15:36:59 +02:00
Adam Osewski	2622ff06cb	Remove unused lds direct load instruction. (#2573 ) This functionality is replaced by amd_async_buffer_load Co-authored-by: Max Podkorytov <4273004+tenpercent@users.noreply.github.com> Co-authored-by: Aviral Goel <aviral.goel@amd.com>	2025-08-06 15:16:12 +02:00
aska-0096	fe63a646a4	add __restrict__ to tr load	2025-08-06 05:58:43 +00:00
Enrico Degregori	2203b0ddfe	Add padding to 1x1Stride1Pad0 conv specialization (grouped conv bwd weight) (#2610 ) * Add padding 1x1Stride1Pad0 conv specialization * Add gridwise checks for conv cshufflev3 * Merge padding with previous transforms * Apply transform changes for padding to default specialization as well --------- Co-authored-by: Bartłomiej Kocot <barkocot@amd.com>	2025-08-05 15:23:19 +02:00
aska-0096	414cad667b	Add XOR fold strategy for hdim<128, but perf dropped; disable it by default; wait further perf debug	2025-08-05 07:23:51 +00:00
Thomas Ning	cbfecf8d7a	Persistent grouped gemm CompV4 Enablement & Polish (#2605 ) * enable the persistent kernel for CompV4 * polish the example and clang format * fix the non-persistent kernel error --------- Co-authored-by: ThomasNing <thomasning@amd.com>	2025-08-04 23:43:01 -07:00
Jinchao Xu	15eb493152	Add -gsplit-dwarf flag to reduce debug section size and fix ckProfiler link errors (#2611 ) Resolves R_X86_64_32 relocation out of range errors in grouped conv2d instances by splitting debug information into separate .dwo files. Add explicit cast to avoid signed/unsigned comparison warning.	2025-08-04 11:26:08 -07:00
aska-0096	0d12fc944f	Add v_permlaneb32 for block_reduce. Disable it as it will cause un-coexecutable packed math in FA	2025-08-04 10:27:42 +00:00
aska-0096	4f31847de1	add vmcnt guard before load ktile	2025-08-04 10:02:17 +00:00
aska-0096	746f4ccb99	Load Q through lds, implement xor;	2025-08-04 06:49:01 +00:00
Illia Silin	788e8a878e	update the switch condition for buffer built-ins (#2602 )	2025-08-01 14:30:07 -07:00
Thomas Ning	7c44a763fa	Fix the GFX 950 Universal GEMM (#2597 ) * solve the gfx950 error * clang format * fix a typo error --------- Co-authored-by: ThomasNing <thomasning@amd.com>	2025-08-01 09:32:24 -07:00
aska-0096	2d4e73d2b4	small refactor	2025-08-01 10:44:54 +00:00
lalala-sh	bb5c478295	fix weight index out of range (#2414 )	2025-08-01 17:50:02 +08:00

1 2 3 4 5 ...

1010 Commits