composable_kernel

mirror of https://github.com/ROCm/composable_kernel.git synced 2026-06-07 00:04:37 +00:00

Author	SHA1	Message	Date
Graner, Johannes	bb6a3571a2	Regression fix for cshuffle	2025-12-05 14:06:53 +00:00
Graner, Johannes	56caf529f8	Code deduplication	2025-12-05 14:06:53 +00:00
Graner, Johannes	785be14874	No more regression, use templates instead	2025-12-05 14:06:53 +00:00
Johannes Graner	cca078ab9c	Merge branch 'develop' into jograner/conv-bwd-weight-kbatch-in-2gb-limit	2025-12-02 08:21:10 +01:00
Erwin Terpstra	46f1d740f0	Add grouped gemm instances for RDNA4 (#3237 ) * wip: grouped_gemm implementation based on wmma kernel + example for fp16 * chore: clean up grouped_gem_wmma_splitk_fp16 example * chore: add cmake options to fully disable XDL or WMMA kernels * feat: add tests for grouped gemma wmma instances for f16 and bf16 (all layouts) * chore: add grouped gemm wmma bf16 example * refactor: reuse more code between instance factory functions * chore: turn test failure if not all batch sizes are supported into a warning * chore: made failing of test on unsupported instances conditional to not break old tests * chore: add log message to failure case where AK1/BK1/KBatch is too high for K value * fix: issue with new overloads of GridwiseGemm_wmma_cshuffle_v3::Run() * fix: stray comma after parameter list * fix: compilation issues on RDNA3 and tests failing due to unsupported problems still being ran * chore: update copyright in header comments * nit: minor feebdack * refactor: unified XDL / wma tests * fix: properly disable FP8 instances when ONLY targeting gfx11 * refactor: add v3 suffix to grouped_gemm device struct name * fix: small typos in example code * fix: fully exclude xdl/wmma instances when using the corresponding cmake flags * chore: remove unused destructor and added pipeline support checks to remove unnecessary paths * fix: make sure to not add instance library to group if library was skipped * fix: make sure xdl grouped gemm doesnt fail the new test * fix: explicitly exclude test if no xdl/wmma support, as pattern matching fails in this case * fix: examples not working since dependent types and functions were moved to ck namespace in develop * fix: tests failing when compiling for just gfx11 due to trying to run unsupported instances * chore: replace/add copyright headers with new format	2025-12-01 15:32:10 -08:00
Graner, Johannes	765fcfc067	update V3 2GB check	2025-12-01 13:18:19 +00:00
Graner, Johannes	2695863e85	Fix variable name in info print	2025-12-01 13:00:31 +00:00
Graner, Johannes	32b3a538d9	Refactor applicability checks into separate function	2025-12-01 12:39:08 +00:00
Graner, Johannes	9e3e1b6935	Fix broken test	2025-12-01 12:38:56 +00:00
Aviral Goel	de6466481f	chore(copyright): update copyright header for include directory (#3293 )	2025-11-26 11:00:05 -07:00
John Shumway	10a782d846	Fix template parameter macros (#3305 ) Some of the device implementation templates have macros like GridwiseGemmMultiABDTemplateParameters that can cause build errors if multiple files are included together. This error comes up with our builder code. To clean up the macros and make them safer, we follow these follow rules: * Use more specific names to avoid duplication. * Undefine the macro after it is used to avoid leaking out of the file scope. * Use a prefix CK_ on the macro to avoid conflicting with other libraries. * Use all caps with underscores for preprocessor macro names.	2025-11-26 09:48:17 -08:00
Graner, Johannes	5d1d298e2b	Disallow hack in non-contiguous edge cases	2025-11-25 13:41:16 +00:00
Graner, Johannes	7ad122d1c1	Clean up comments	2025-11-24 09:40:50 -05:00
Graner, Johannes	0e8ca436f3	Disable hack if stride not divisible by k_batch	2025-11-24 09:19:30 -05:00
Graner, Johannes	350162728d	Fix buffer size calculations	2025-11-24 09:19:30 -05:00
Graner, Johannes	fc86ec44f5	Don't miss half of the elements	2025-11-24 07:25:12 -05:00
Graner, Johannes	e63bc7a2ce	Fix tensor descriptors and stride calculations	2025-11-24 07:24:06 -05:00
lalala-sh	f58bd56e6b	fix static assert (#3178 ) Co-authored-by: Illia Silin <98187287+illsilin@users.noreply.github.com>	2025-11-20 17:27:05 -08:00
Gavin Zhao	07314ac543	Add support for RDNA1 GPUs (#3220 ) * Allow compilation for RDNA1 (__gfx101__) Signed-off-by: Gavin Zhao <git@gzgz.dev> * More RDNA1 changes Signed-off-by: Gavin Zhao <git@gzgz.dev> * Even more RDNA1 changes Signed-off-by: Gavin Zhao <git@gzgz.dev> * cmake: skip build quantization for unsupported arches * add gfx10-1-generic support as well * add gfx1013 and complete gfx10-1-generic * fix clang format * enable DL kernels on gfx101x --------- Signed-off-by: Gavin Zhao <git@gzgz.dev> Co-authored-by: illsilin_amdeng <Illia.Silin@amd.com> Co-authored-by: Illia Silin <98187287+illsilin@users.noreply.github.com>	2025-11-20 10:45:57 -08:00
Aviral Goel	f5ac3ee359	chore(copyright): update copyright header for include directory (#3224 ) * chore(copyright): update copyright header for tile_engine directory * chore(copyright): update copyright header for script directory * chore(copyright): update copyright header for test_data directory * chore(copyright): update copyright header for python directory * chore(copyright): update copyright header for profiler directory * chore(copyright): update copyright header for library directory * chore(copyright): update copyright header for include directory	2025-11-18 10:17:18 -08:00
jefyang1	d30babbd00	Add new gemm multiply multiply instances on gfx950 (#3213 )	2025-11-14 08:20:41 -08:00
yinglu	2a73eb3bc0	Simulate TF32 with BF16x3 (#3142 ) * tf32:bf16x3:use bf16x3 emulate tf32 gemm * change blockwiseGemm to demo bf16x3 * temp push * self review * self review * fix multi-device compile error * bug fix * code refactor * limit to gfx950 * enhance gemm gfx942 threshold * lower change from blockwise to warpwise * refact codes * refact codes * error fix * change threshold * bug fix * fix threshold error * change host reference implement to same as device * bug fix * bug fix * code refact * fix clang-format fail * code refine	2025-11-13 16:21:09 -08:00
Enrico Degregori	7414a0f4d4	Wmma support for gemm_reduce (#3145 ) * Initial implementation GEMM+Reduce: - device struct - epilogue struct * Fix tests, improve profiler and add initial instances * Add instances * Fix compilation error * Address review comments * Fix logging --------- Co-authored-by: Illia Silin <98187287+illsilin@users.noreply.github.com>	2025-11-12 11:23:54 -08:00
Enrico Degregori	1c544abf57	Extend support for ak1 / bk1 WMMA (#3073 ) * Extend AK1 / BK1 support: - Add support for AK1 != BK1 - Add support for AK1, BK1 > 8 - Introduce KInner template parameter for pipelines when loading multiple tiles with one instruction * fix clang format	2025-11-11 07:38:15 -08:00
Bartłomiej Kocot	64ca8414d6	Update gridwise_gemm_xdl_cshuffle_conv_v3.hpp (cherry picked from commit `b33877c10f`)	2025-11-07 13:15:33 +00:00
Bartlomiej Kocot	59159f1962	Optimize grouped conv bwd wei split_k off calc (cherry picked from commit `6f61dd56c5`)	2025-11-07 13:15:14 +00:00
Graner, Johannes	4fa6869f81	Revert "Take split_k into account when checking 2GB tensor limit." This reverts commit `adf35c91be`.	2025-11-07 13:11:51 +00:00
Graner, Johannes	b9f4c84170	Merge branch 'develop' into jograner/conv-bwd-weight-kbatch-in-2gb-limit	2025-11-07 12:23:45 +00:00
Gino Lu	e31a7a4f29	fix MX bpreshuffle gemm B grid descriptor dimension error. (#3170 )	2025-11-06 19:42:39 -08:00
Xudong Yuan	d04eba4ae3	Ck moe mxfp4 blockm32 (#3098 ) * block_m = 32 * ck block_m = 32 * aiter/3rdparty/composable_kernel/include/ck/tensor_operation/gpu/block/blockwise_gemm_pipeline_xdlops_b_preshuffle_mx_moe_v3.hpp format * mxfp4_moe v1 pipe * update format --------- Co-authored-by: zhimding <zhimding@amd.com> Co-authored-by: lalala-sh <Jiaxing.Wen@amd.com> Co-authored-by: felix <felix.li@amd.com>	2025-11-07 08:45:41 +08:00
Graner, Johannes	adf35c91be	Take split_k into account when checking 2GB tensor limit.	2025-11-06 09:45:29 +00:00
Adam Osewski	b8527a9236	[CK_BUILDER] Convolution traits. (#3152 ) Added: 1. Convolution traits & unit tests 2. Update builder enumerators to have representation of Convolution Kernels properties. 3. Unified builder pipeline version & scheduler enumerators	2025-11-05 08:53:06 -08:00
Illia Silin	930423ab3b	Initialize new variable to prevent c++17 compiler error (#3156 ) * initialize new variable to prevent c++17 compiler error * build for gfx90a using -std=c++17 flag	2025-11-04 18:54:14 -08:00
John Shumway	6dbee64886	[CK_BUILDER] Add backward weight instance traits for xdl cshuffle. (#3143 ) * Add backward weight instance traits for xdl cshuffle. To keep instance test file sizes reasonable, we start a new test_bwd_weight_instances_traits.cpp test file. * Fix copyright notices. * Remove (c) symbol, replace with (C). Having UTF-8 in source caused an error with code generation.	2025-11-04 15:34:00 +01:00
Enrico Degregori	507d81c3af	Fix splitk preshuffle (#3137 ) * Fix splitK multiply_multiply_wp * Add tests for gemm_multiply_multiply_wp * Add tests for gemm_universal_preshuffle (KBatch = 1) * Add tests gemm_blockscale_wp * Fix splitk gemm universal preshuffle * Run new tests on arch supporting fp8 * Restore example * Fix strides profiler * Fix tests * Fix clang format * Finalize profiler preshuffle with tolerances * Minor improvements to splitk related changes * Address review comments: clang format and ckProfiler typo * Remove b_k_split_offset from SplitKBatchOffset struct	2025-11-03 11:59:01 -08:00
Bartłomiej Kocot	ab1a8356b6	Add 2GB limitation for grouped conv bwd weight (#3054 )	2025-11-01 14:16:45 +01:00
Enrico Degregori	4ebc48a3cd	WMMA gemm_add_relu_add_layernorm (#2989 ) * Summary: - Refactor epilogue (with CShuffle) to support fused operations: - EpilogueCShuffleBase holds common parts - EpilogueCShuffle: runs CShuffle and write out - EpilogueWelfordCShuffle: holds Welford specific arguments, runs CShuffle, write out, Welford first part and Welford write out - Extend thread transfer v7r3: - Support for intermediate data type different from src and dst type - New functionality to write to dst buffer and keep data (to be able to use them for additional operations) * Adress review comments	2025-10-31 11:19:26 -07:00
John Shumway	5ed2046bee	Add the last two forward instance traits. (#3134 ) * Add InstanceTraits for DeviceGroupedConvFwdMultipleD_Wmma_CShuffle * Add InstanceTraits for kernel_grouped_conv_fwd_dl_multiple_d * A few small changes to fix broken instance traits.	2025-10-31 07:52:42 -07:00
kabrahamAMD	a7c52e8afa	Kabraham/fix block gemm v1 b scale (#3129 ) * fixed synchronization issue in block gemm pipeline v1 that caused b_scale to fail * run clang-format --------- Co-authored-by: Kevin Abraham <kevin.abraham@streamhpc.com>	2025-10-31 07:19:01 -07:00
John Shumway	cafaeb6b7b	Add instance traits for two more grouped forward convolutions (#3112 )	2025-10-29 16:04:13 +01:00
Bartłomiej Kocot	66bae4306c	Grouped conv fwd with direct load (#3082 ) * Grouped conv fwd with direct load * fix * fix * Add IsSupported check * Fix * fix inductor	2025-10-29 09:54:42 +01:00
Ville Pietilä	1c17bae816	Add name member to CK elementwise operations. (#3102 )	2025-10-27 22:19:29 -07:00
John Shumway	54746e9329	[CK_BUILDER] Test and fix instance traits utils. (#3096 ) * Refactor instance_traits_util and add unit tests tests * Address reviewer comments. Just adds some TODOs to indicate deprecated layouts in our reflection. Our strategy is to leave the reflection code broad (covering deprecated features), but keep the builder concepts narrow. Once we've removed deprecated features from all instances, we can remove them from reflection. Also add a comment to the cmake to explain the unit test target test_conv_builder. * Addressed more reviewer comments. * Remove duplicate PassThrough::name Accidentally added this field to the end of the struct, too. The `name` field should be a the start of the struct for consistency.	2025-10-27 22:14:08 -07:00
Ville Pietilä	6c2ca1211a	[CK_BUILDER] First fwd convolution builder implementation (#3070 ) * Add experimental builder infrastructure for composable_kernel - Add experimental/builder directory with README documentation. - Create initial test infrastructure with CMakeLists.txt and placeholder test. - Update root CMakeLists.txt to support CK_EXPERIMENTAL_BUILDER option. - Update .gitignore to not treat `experimental/builder` as a CMake build directory. This establishes the directory structure for a high-level builder pattern that will provide a semantically-clear interface for constructing CK operations, with initial focus on convolution kernels for MIOpen integration. * Fix clang formatting. * Fix CMake build infrastructure for experimental builder - Add experimental/builder CMakeLists.txt with proper subdirectory structure - Add placeholder include/ck_tile/builder CMakeLists.txt for header installation - Fix gtest.cmake to use include_guard to prevent multiple inclusions - Update root CMakeLists.txt to include full builder directory instead of just tests * Scope C++20 settingto the test code Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com> * Remove redundant GTest::gtest linkage Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com> * Introduce basic types, and convolution algorithm concepts and limits. * Add convolution signature concepts. * Add convolution factory. * Finalize conv factory implementation for fwd convolutions. * Add type definitions for testing. * Add placeholder test. * Add convolution builder definition. * Fully functional fwd conv builder. * Test improvements. * Clean-up include headers. * Enable the limit checks for the convolution algorithm parameters. * Remove dead code. * clang formatting. * Add more tests and missing conv specialization argument. * clang formatting. * Add explicit handling of the tensor layouts. * Add complete 2D/3D layout support to CK Builder - Add missing 2D layouts: GNHWC_GKYXC_GNHWK, NGCHW_GKCYX_NGKHW - Add missing 3D layout: GNDHWC_GKZYXC_GNDHWK - Add 1D layouts (NWGC, NGCW, GNWC, NGCW_GKCX) for future support - Add 3 tests for new 2D/3D layouts - All tests pass (5/5) * Add tests for remaining 2D/3D layouts - Add test for 2D NGCHW_GKYXC_NGKHW (channels-first) with Filter1x1Stride1Pad0 - Add test for 3D NDHWGC_GKZYXC_NDHWGK (channels-last) - All 7 tests pass (complete coverage for all 2D/3D forward layouts) * Change enum converters to consteval. * 7 tests with pipeline and specialization\| Test # \| Dim \| Type \| Layout \| Pipeline \| Specialization \| \|--------\|-----\|------\|----------------------\|----------\|-------------------------\| \| 1 \| 2D \| BF16 \| NHWGC_GKYXC_NHWGK \| V1 \| DEFAULT \| \| 2 \| 2D \| FP16 \| GNHWC_GKYXC_GNHWK \| V3 \| FILTER_1X1_PAD0 \| \| 3 \| 2D \| FP32 \| NGCHW_GKCYX_NGKHW \| V4 \| FILTER_1X1_STRIDE1_PAD0 \| \| 4 \| 2D \| BF16 \| NHWGC_GKYXC_NHWGK \| V5 \| FILTER_3x3 \| \| 5 \| 3D \| FP32 \| NGCDHW_GKCZYX_NGKDHW \| V1 \| FILTER_1X1_PAD0 \| \| 6 \| 3D \| BF16 \| GNDHWC_GKZYXC_GNDHWK \| V3 \| DEFAULT \| \| 7 \| 3D \| FP16 \| NDHWGC_GKZYXC_NDHWGK \| V4 \| FILTER_1X1_PAD0 \| * Add missing convolution layouts and provide better compile-time error in instance traits. * Fix clang formatting. * Changed I8 -> S8. * Fix signature. * Rename concepts and corresponding members. * Rename LDS related parameters. * Remove ODD_C specialization. Add V2 pipeline. * Add missing types. * Add elementwise operation to the conv signature. * Improve compile-time error message for unsupported elementwise ops. * Separate different fwd conv builder tests into separate compilation units. * Fix layout to string and add name to old CK PassThrough elementwise op. * Enable both CK and CK Tile tensor layouts in instance traits. * Fix clang-format. --------- Co-authored-by: John Shumway <jshumway@amd.com> Co-authored-by: John Shumway <john.shumwayjr@gmail.com> Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com> Co-authored-by: JH-Leon-KIM-AMD <jeonghyun.kim@amd.com>	2025-10-27 20:09:24 +02:00
yinglu	6bbc05e1bd	conv:tf32:add missed instances (#3081 ) * conv:tf32:add missed instances	2025-10-24 16:28:36 +08:00
John Shumway	37dff024c1	[CK_BUILDER] Add compile-time reflection for a convolution instance (#3065 ) * [CK_BILDER] Add compile-time reflection for a convolution instance Introduce InstanceTraits template metaprogramming framework to enable runtime introspection of device kernel template parameters without requiring implementation knowledge. This reflection system extracts configuration details (block sizes, data types, layouts, tuning parameters) directly from kernel specializations through template pattern matching. In particular, the GetInstanceString method returns a string that uniquely idenitfies the kernel, by explicitly serializing all template paramter values. This provides critical functionality for MIOpen integration, since the existing GetTypeString method is ambiguous, and only captures some of the template paramters. The implementation uses a two-level design: a primary InstanceTraits template declaration in instance_traits.hpp serves as the interface, while kernel-specific specializations (e.g., for DeviceGroupedConvFwdMultipleABD_Xdl_CShuffle_V3) provide the actual extraction logic. This separation allows the reflection system to scale to additional kernel types without modifying the core interface. Key architectural decisions: - Forward-declare device kernels in instance_traits.hpp to avoid circular dependencies, since device implementation headers will include the reflection headers - Use compile-time constants and type aliases to expose kernel parameters, enabling zero-overhead introspection - Provide a templated instance_string() function that generates human-readable kernel configuration strings by serializing all template parameters in order, useful for debugging and kernel identification - Guard reflection integration with preprocessor definition CK_EXPERIMENTAL_BUILDER to keep it opt-in until the API stabilizes - Add GetInstanceString() virtual method to BaseOperator, allowing runtime polymorphic access to compile-time kernel information This infrastructure also enables upcoming higher-level semantic reflection abstractions (like ConvTraits) to query kernel configurations programmatically. Includes unit tests validating both the trait extraction accuracy and the string generation format.	2025-10-21 21:10:19 -07:00
Bartłomiej Kocot	3a28632b20	Gridwise gemm conv v3 force padded layout on gfx950 (#2961 ) * Gridwise gemm conv v3 force padded layout on gfx950 * fix bug in other gridwise * fix * Update gridwise_gemm_wmma_cshuffle_v3_common.hpp	2025-10-21 15:41:02 +02:00
Ville Pietilä	7e44b845b5	Fixed handling of split-K autodeduce argument for grouped convolution (#3024 ) * Fix handling of split-K autodeduce argument. * Fix clang formatting. * Test fix. * Fix clang formatting.	2025-10-17 15:36:39 +03:00
Enrico Degregori	440358c168	Wave Tile Transfer supporting global load with transpose (#3027 ) * Initial implementation: - add new thread group transfer supporting transpose instruction - refactor AB transfer to switch between thread and wave tiles methods * Add some comments and remove explicit wave and lane calculations * Remove compiler option for performance * fp16 example: use tuned instance * Missing cleanup * Integrate wave transfer in existing gemm and batched gemm instances * Add fast instances * extend implementation for 8 bit datatypes packed types not supported * Address review comments * Optimize pipeline v1 and re-introduce compiler option * Disable wave tile approach for b scale gemm * Fix for clang20 * Avoid code duplication of amd_global_load_transpose_to_vgpr function	2025-10-16 11:33:56 -07:00
kabrahamAMD	c4b2da9cbd	implement device batched gemm b scale for wmma (#2825 ) * rebased on top of develop * fixed missing shuffeling and wrong indexing * added tests for batched_b_scale * added missing files * fixed wrong stride computation and removed k batching (for now) due to precision issues * reinstated k-batching with PRNG constrained to -1..1 * added specialization of GeneratorTensor_3 for int4 and fixed internal overflow * added k-batching to reference and increased tolerances for test * changed gemm_b_scale and gemm_universal tests to use correct parameters * adressed review commentsd * ported fixes back to non-batched version of b_scale * adressed review comments * run clang-format on older commits * add type-conversion to AccDataType and then to CDataType to exactly mimic GPU's behavior * added newline at end of file * reflected changes from muitl-abd branch in batched b_scale * fixed gfx11 issue * changed range for pki4 to -1...1 (-0.5...0.5 never really made sense for i4 anyway and always should have caused compiler errors, but since there was no int4 specialization of GeneratorTensor3 until now, this passed * run clang format * set range of i4 generation to 0...1 for upstream tests to pass. This replicated previous behavior, which however means that it is NOT properly tested. * reduced range for pk_i4 even further to 0..0 * removed failing xld instances. Failure now uncovered now that tests were fixed * removed generation of int4 values entierly * divide B buffer by BPackedSize --------- Co-authored-by: Kevin Abraham <kevin.abraham@streamhpc.com>	2025-10-16 11:00:42 -07:00

1 2 3 4 5 ...

619 Commits