composable_kernel

mirror of https://github.com/ROCm/composable_kernel.git synced 2026-05-11 17:00:18 +00:00

Author	SHA1	Message	Date
Illia Silin	1b7da171c9	Update the list of contributors. (#836 ) * add linting and update contributors list * skip the linting and doc changes * add Astha * add YanXing	2023-08-09 13:44:13 -07:00
Illia Silin	9af519ee86	add gfx941 to the ckProfiler package (#840 )	2023-08-09 10:30:40 -07:00
Bartłomiej Kocot	472fa029ba	Enable grouped conv with small K or C (#822 ) * Enable grouped conv with small K or C * Add missing instances * Refactor grouped conv fwd instances * Fix fp16 instances since it supports src_per_vec %2 = 0 * Add generic instances	2023-08-09 10:40:55 -05:00
Rostyslav Geyyer	9c54eaab04	Enable f16/f8 mixed precision mode (#820 ) * Enable f16/f8 mixed precision * Add an argument to enable mixed precision * Update for compatibility * Add mixed precision example * Introduce ComputeType argument	2023-08-09 08:44:23 -05:00
Illia Silin	6802611334	add no-offload-uniform-block flag for rocm5.7 and up (#838 ) * add -fno-offload-uniform-block flag for rocm5.7 and up * add a comment and compiler ticket number	2023-08-08 17:58:31 -07:00
Illia Silin	08eb176929	Allow building CK for specific data types and split off last remaining DL instances. (#830 ) * properly split conv_nd_bwd_data instances * split conv2d_fwd instance data types * split the gemm, conv2d_fwd and batched_gemm_softamx_gemm * split the tests by data types where possible * filter examples by DTYPES * split few remaining examples by DTYPES * filter most instances by DTYPES * add new lines at end of headers, fix grouped_gemm profiler * fix syntax * split the ckprofiler instances by DTYPES * split the conv2d and quantization DL and XDL instances * fix the splitting of conv2d DL instances * split softmax and pool_fwd tests for fp16 and fp32 types * fix syntax * fix the dl_int8 quantization instances isolation	2023-08-07 14:56:10 -07:00
Bartłomiej Kocot	22443f7aae	Add wei_strides to grouped conv3d wei to keep consistency (#817 ) * Add wei_strides to grouped conv3d wei to keep consistency * Fix strides in client examples * Unify backward weight api with forward * Fix for example * Fixes for examples --------- Co-authored-by: zjing14 <zhangjing14@gmail.com>	2023-08-07 10:23:45 -05:00
Illia Silin	2474dddbee	add an option to build ckProfiler package for specific architectures (#828 )	2023-08-03 10:10:27 -07:00
Bartlomiej Kocot	aac65a031e	Change to github_issue prefix	2023-08-03 16:38:28 +02:00
Bartlomiej Kocot	e6a826d35a	Rename the workaround to a proper issue name	2023-08-03 16:38:28 +02:00
Bartlomiej Wroblewski	8c13df07bf	Improve formatting of docs; Add a note about the DL_KERNELS flag (#825 ) * Improve formatting of docs; Add a note about the DL_KERNELS flag * Change the recommended version of ROCm to 5.6	2023-08-03 15:50:38 +02:00
Po Yen Chen	f7cc8c3b03	Update tuning parameter & compilation options of DeviceGemmXdl<> instance (layout=TT) (#819 ) * Enable pipeline v2 opt for layout=TT instance * Use better thread mapping for reading A tile * Conditionally enable pipeline v2 opt * Allow enabling only fp16 gemm instances in profiler * Fix formatting error * Fix compilation error if we enable fp32 in profiler	2023-08-02 10:32:22 -05:00
Bartłomiej Kocot	7761e5232c	Add s_nops after v_dot to avoid hazard (#808 ) * Add s_nops after v_dot to avoid hazard * Fix builtin for inner_produxt fp16 * Skip inline version to builtin * Add comments regarding isa * Fix comment regarding s_nop	2023-07-27 13:29:44 -05:00
carlushuang	e7dca79d27	initial stream-k implementation with example (#699 ) * initial stream-k implementation with example * fix unexpected change in err * improve a little bit performance by reorganize pipeline. * improve perf a little bit by swizzle block idx * add profiler * update example * fix spelling * shrink karg for streamk * support dynamic buffer using memory coherence glc_slc bit from template * control memory coherence while construct dynamic buffer * update reduction for streamk(not ready yet) * Add template parameter to make_dynamic_buffer to support amd_buffer coherence setting * fix build issue * fix several bug * now result is correct, everything works (but has scratch) * remove scratch by manually reset coordinate * update device code * fix a bug in final reduce * fix something in example * update async memset * fix enum as camel case * modify coherence enum name * clean code and use atomic streamk by default * remove unused var * throw exception if have empty pointer * fix format * fix CI warning * fix type in init * modify CI error * filter out on gfx10+ * restore changed example code --------- Co-authored-by: Qianfeng Zhang <Qianfeng.Zhang@amd.com>	2023-07-26 14:18:15 -05:00
Illia Silin	9195435c77	Disable DL kernels by default. (#816 )	2023-07-26 11:06:45 -05:00
Bartłomiej Kocot	ac6d68b353	Disable XDL kernels on unsupported HW Add ck::is_xdl_supported (#768 ) * Disable XDL kernels on unsupported HW; Add ck::is_xdl_supported function (#765) * Do not throw an error when GEMM problem is not supported. --------- Co-authored-by: Bartlomiej Wroblewski <bwroblewski10@gmail.com> Co-authored-by: Adam Osewski <aosewski@amd.com> Co-authored-by: Illia Silin <98187287+illsilin@users.noreply.github.com>	2023-07-26 07:19:55 -07:00
rocking	016bd428df	Refine the dimension of host tesnor. This example only require 1D (#812 )	2023-07-25 23:18:56 -05:00
Po Yen Chen	f4ea560112	Speed-up global memory reading for GEMM instances (#813 ) * Use better ThreadClusterLengths to speed up * Update B tile reading pattern for layout=NN instance	2023-07-25 18:54:47 -05:00
ltqin	50643dd555	Add bias scalar vectorload = 1 for gemm bias gemm (#791 ) * first change bias load * add bias dim and scalervector parameter * make CDE0BlockTransferSrcVectorDim not work * changse toinstance * add limit for CDE0BlockTransferSrcScalarPerVector	2023-07-24 20:08:15 -05:00
Illia Silin	844b215d92	add ninja profiling tools to the base docker (#805 )	2023-07-21 15:33:17 -07:00
Illia Silin	7a29f711d4	add INSTANCES_ONLY cmake macro to build only instances (#807 )	2023-07-21 15:31:19 -07:00
Bartłomiej Kocot	10732847e7	Grouped conv bwd wei NDHWGC/NDHWGK (#804 )	2023-07-21 12:00:55 -05:00
Bartłomiej Kocot	49180fd60b	Grouped 3d conv backward data support (#799 ) * Grouped 3d conv backward data support * Fix comments	2023-07-18 11:01:33 -05:00
Rostyslav Geyyer	f82bd59389	Remove type_convert bf16 to int32 and back (#802 )	2023-07-18 09:44:51 -05:00
Illia Silin	189ea3b9aa	Add mechanism to build CK for select data types, add Navi3x CI. (#790 ) * allow building CK for specific data types * add CI build and test stage on Naiv3x without some int8 instances * add missing gemm fp16 instances * add the changes to the missed cmake file * add empty lines at end of source files * Do not build quantization client example on navi3 in CI * disable batched_gemm_multi_d_int8 instances with DTYPES * disable device_conv2d_bwd_data_instance with DTYPES * fix ckprofiler for conv_bwd_data for int8 * properly isolate the conv_bwd_data int8 instances * remove empty line	2023-07-17 18:02:42 -07:00
Illia Silin	4867db4290	Add check for compiler GPU target support. (#800 ) * check if gpu_targets are supported by compiler * set default list of targets and filter for them	2023-07-17 09:44:40 -07:00
arvindcheru	03d3395b3c	Disable Werror to ignore xnack+ warnings (#794 ) * Disable Werror to ignore xnack+ warnings	2023-07-14 20:00:20 -04:00
Bartłomiej Kocot	1ee99dcaa6	Support NHWGC conv2d_bwd_weight (#769 ) * Support NHWGC conv2d_bwd_weight * Fix client example * Fix client example * Fix comments * Redesign grouped_conv_bwd_weight instances * Clang format fix --------- Co-authored-by: zjing14 <zhangjing14@gmail.com>	2023-07-12 08:25:02 -05:00
Illia Silin	87f2bbcf5c	change the build thread usage in CI (#787 )	2023-07-06 20:17:25 -05:00
Adam Osewski	237f9cd3aa	Add basic setup for precommit (#749 ) (#764 ) * Add basic setup for precommit * Update README.md with instructions on installing precommit hooks --------- Co-authored-by: Illia Silin <98187287+illsilin@users.noreply.github.com> Co-authored-by: Bartlomiej Wroblewski <bwroblewski10@gmail.com>	2023-07-06 11:01:06 -05:00
Po Yen Chen	850144a0d3	Split GEMM instance library & enable pipeline v2 optimization (#783 ) * Move source file into sub-directories * Add missing include directive * Split DeviceGemmXdl<> fp16 instances * Fix format * Remove unnecessary CMakeLists.txt * Add macros to toggle new features * Remove debug message * Turn off GEMM v2 pipeline optimization by default * Fix format * Extract duplicated string as list * Enlarge indent in CMakeLists.txt	2023-07-06 10:59:35 -05:00
Qianfeng	8f5cafaf04	Batchnorm splitk single kernel (#771 ) * Use dim 0 as faster dim for writing mean/var/count workspace in batchnorm multiblock method [performance] * Add CountDataType as template parameter in blockwise_welford * Add utility/get_shift.hpp * Add BatchNorm multiblock single-kernel implementation * Add smem inline assembly based implementation of gms_init/gms_barrier/gms_reset for gfx90a * Renaming in device_batchnorm_forward_impl.hpp * Tiny fix in the batchnorm_fwd profiler * Revert "Add smem inline assembly based implementation of gms_init/gms_barrier/gms_reset for gfx90a" This reverts commit `d16d00919c`. * Use the old two-kernel batchnorm multiblock method for gfx1030 * Use the old two-kernel batchnorm multiblock method for gfx908 * use the single-kernel batchnorm multiblock method only for gfx90a * Remove get_wave_id() from utility/get_id.hpp since it is not used * Set true for testing running mean/variance and saving mean/invvariance in the examples * Fix to copy-right words * Remove un-needed including in utility/get_id.hpp * Add comments to workgroup_synchronization.hpp * Remove un-used codes in gridwise_multiblock_batchnorm_forward.hpp * Renaming in the kernels * Remove un-used kernel file	2023-07-06 10:58:55 -05:00
Adam Osewski	f4dfc060b7	Move Device Ops implementations into impl directory. (#777 ) Co-authored-by: Adam Osewski <aosewski@amd.com> Co-authored-by: zjing14 <zhangjing14@gmail.com>	2023-07-06 16:15:51 +02:00
Bartlomiej Kocot	2b0b6d9f46	Fix copyrights for DeviceBatchedGemmMultipleD_Dl	2023-07-06 15:50:27 +02:00
Rostyslav Geyyer	61dc9aa932	Add the missing archs (#785 )	2023-07-05 18:29:56 -05:00
Rostyslav Geyyer	1cf5003179	Add fp8 GEMM and an example for it (#767 ) * Add fp8 xdl gemm * Add example * Use int8 intrinsics for buffer load/store * Format * Update cmakelists	2023-07-04 20:38:49 -06:00
Illia Silin	7797bd3d2b	Upgrade default docker to ROCM5.6 release. (#778 ) * upgrade default compiler to rocm5.6 release * do daily runs with rocm5.6 instead of 5.5	2023-06-30 08:06:54 -07:00
Illia Silin	d3adc66581	Add rocm5.6 RC4 and rocm5.7 to docker build options. (#770 ) * upgrade to rocm5.6 rc4 * add rocm5.7 docker	2023-06-28 08:58:28 -05:00
Illia Silin	3b18f1e38c	do not build gfx941/942 targets during CI (#766 )	2023-06-21 10:47:35 -07:00
Bartłomiej Kocot	63388e84ab	Support bf16/f32/f16 and NHWGC conv2d_bwd_data (#757 ) * Support bf16/f32/f16 and NHWGC conv2d_bwd_data * Add interface test * clang format * Comment fixes * Add more friendly error message	2023-06-21 08:20:31 -05:00
ltqin	32d2f52bf7	remove useless comments (#760 )	2023-06-19 19:25:08 -07:00
zjing14	05ea6452b6	changed pipeline v1 (#763 )	2023-06-19 19:24:18 -07:00
Illia Silin	645eb2f2a0	do not build gemm-gemm and conv-conv examples for gfx94* (#761 ) * do not build gemm-gemm and conv-conv examples for gfx94* * do not build gemm-gemm and conv-conv examples on navi	2023-06-19 16:55:03 -07:00
Rostyslav Geyyer	f0c620c42e	FP8 enablement - add a pseudorandom number generator, add conversion methods (#708 ) * Add basic fp8 definitions and prn-generator * Format * Add fp8<->fp32 type_convert * Format * Split type_convert and cast_to/from_f8 * Format * Minor fix * Minor fix * Move fp8 utils to a separate header * Add elementwise ops * Add fp8_convert_sr * Format * Add element op * Eliminate magic numbers * Split f8_convert_sr in host and device * Format * Add some constexpr * Add a datatype test * Format * Another format * Add fp8<->fp16 tests * Update type_converts * Format * Add fp16 casting functions * Format * Use seed as a runtime arg * Use element location for PRNG * Format * Add fp8<->fp16 to PassThrough element op * Clean up * Merge host and device implementations * Add comments on rounding modes * Remove leftover code * Put type_converts into a separate header * Put random number gen to a separate header * Rearrange f8_utils' namespaces * Refactor type_convert.hpp * Move f8_t definition	2023-06-19 11:20:35 -05:00
rocking	341ad95665	Maxpool bwd (#750 ) * Add maxpool f32 kernel and example * Revise copyright * Add device pool bwd device op * Support f16 and bf16 * Add compute datatype for reference code. Prevent error in bf16 * Fix type error * Remove layout * Fix bf16 error * Add f16 and bf16 example * Add more operations * Implement IsSupportedArgument * Add changelog * Add comment * Add comment * Remove useless header * Move initialize of workspace to the run * Move set din zero to the device operator * Save din_length_raw * Remove useless header * Calculate gridsize according to the number of CU * Calculate gridSize according to the number of CU. Remove useless header * Add put example * Remove useless header * Fix CI fail	2023-06-19 09:44:22 -05:00
Qianfeng	0d9118226b	Padded Generic Kernel Instance (#730 ) * Add NumReduceDim template parameter to DeviceSoftmax and Softmax client API to simplify instances collecting * Move the generic kernel instance to be the first of the instance list for elementwise op of normalization * Add GetGenericInstance() interface for DeviceOperationInstanceFactory class of DeviceSoftmax * Add testing of GetGenericInstance() in client_example of Softmax * Revert "Add testing of GetGenericInstance() in client_example of Softmax" This reverts commit `f629cd9a93`. * Revert "Add GetGenericInstance() interface for DeviceOperationInstanceFactory class of DeviceSoftmax" This reverts commit `a9f0d000eb`. * Support generic kernel instance to be the first instance returned by GetInstances() for GroupNorm * Move generic kernel instance to separate tuple for elementwise op of normalization * Remove un-used files for softmax instance * Store generic kernel instance to separate tuple for softmax * Add IsSupported checking for generic instance to client example of softmax * Replace the get_device_normalize_from_mean_meansquare_instances() by the DeviceOperationInstanceFactory class for elementwise-normalization * clang-format fix * Remove int8 from softmax instances --------- Co-authored-by: zjing14 <zhangjing14@gmail.com>	2023-06-16 23:43:11 -05:00
Illia Silin	d140bdc9fa	do not build gfx941/942 targets during daily QA runs (#758 )	2023-06-16 12:13:16 -07:00
Illia Silin	027e46ee82	Enable gfx941 and gfx942 architectures. (#752 ) * enable gfx941/942 targets * fix clang format * fix the cmake logic for multiple targets * fix cmake syntax for looping over targets * add gfx941/942 support for gemm_xdl instances	2023-06-15 08:20:59 -07:00
zjing14	309b1c6461	Fixed Weight layout of grouped_conv 3d fwd (#743 ) * Changed wei layout * changed layout for examples * fixed client example --------- Co-authored-by: root <root@ctr-ubbsmc15.amd.com>	2023-06-15 10:19:33 -05:00
Qianfeng	c5f6ec842c	Using number of compute units to set gridSize (#754 ) * Add getAvailableComputeUnitCount() interface * Use available number of compute units to set kernel grid size	2023-06-15 10:13:59 -05:00

1 2 3 4 5 ...

955 Commits