composable_kernel

mirror of https://github.com/ROCm/composable_kernel.git synced 2026-05-11 17:00:18 +00:00

Author	SHA1	Message	Date
Illia Silin	180e572076	Fixing most of the cppcheck errors. (#1142 ) * fix cppcheck errors, first pass * fix format * fix returned value in examples * add macro definitions for cppcheck * fix the profile_gemm logic * update the gemm profiler logic * add more difinitions to cppcheck, fix couple more errors * replace runtime error with message in device function * fix a couple of int4 issues * no return for fill function * fix errors in data_types.hpp * fix format * fix few remaining errors * fix errors in data_types.hpp * fix last couple of errors in datat_types.hpp	2024-01-24 13:47:48 -08:00
Bartłomiej Kocot	7e4eb4b800	Add optimized copy to ck wrapper (#1126 ) * Add optimized copy to ck wrapper * Example optimizations * Fixes * Move img2col test to client example * Refactor example * Fix docs * Fixes * Fix * Fixes * Fixes * Fixes * Fixes * Fixes --------- Co-authored-by: zjing14 <zhangjing14@gmail.com>	2024-01-19 11:29:00 +01:00
Bartłomiej Kocot	4234b3a691	Add tensor partition and generic copy for ck wrapper (#1108 ) * Add tensor partition and generic copy for ck wrapper * Update changelog * Stylistic fixes * Change shape/strides logic to descriptor transforms * Fixes * Fix client example * Fix comments	2024-01-03 01:10:57 +01:00
Rostyslav Geyyer	6891e4d109	Fix the bugs (#1099 )	2023-12-13 12:27:31 -08:00
Bartłomiej Kocot	836b7e557d	Introduce wrapper library (#1071 ) * Introduce wrapper library * Update cmake files * Revert "Update cmake files" This reverts commit `c27f88b565`. * Fix comments	2023-12-06 11:58:59 +01:00
Bartlomiej Wroblewski	bc4bf9bd03	Add support for double buffering in direct load GEMM kernel (#1052 ) This PR introduces support for double buffering in LDS into GEMM kernels that use direct load instructions. Direct loads now use inline asm instead of intrinsics. Usage of intrinsics results in compiler adding additional waitcnt instructions what breaks possible load/compute overlap in case of double buffering. Usage of inline asm results in the need to use sched_barrier in order to make sure that compiler cannot incorrectly reschedule instructions since it does not know the data dependencies between global->LDS and LDS->registers.	2023-12-03 23:08:47 +01:00
Bartłomiej Kocot	8ff845f2c4	Introduce wrapper for layout (#1054 ) * Introduce wrapper for layout * Extend functionality * Fix for getLength * Comment fixes * Add comments and remove not needed getters	2023-11-30 12:11:43 +01:00
Rostyslav Geyyer	6ef034f6ca	Switch default f8 conversion to stochastic rounding (#1048 ) * Switch default f8 conversion to stochastic rounding * Refactor f8-related type_converts * Add an element-wise op	2023-11-27 20:06:17 -06:00
Bartlomiej Wroblewski	627054b941	Add basic support for direct loads from global to LDS (#999 ) * Add basic support for direct loads from global to LDS * Clean the code and comments * Add support for fp16 * Add comments * Add check for thread cluster lengths * Align non-direct-load fp16 example * Small fixes * Extend IsSupported to check for supported GPU gens * Build examples only on the supported HW * Do not throw when instance not supported in 04 example * Review: Apply review suggestions * Review: small fix * Review: small fix	2023-11-25 13:35:22 +01:00
zjing14	98fd41f597	Add Gemm instances for performance improvement (#1018 ) * improve kpad * more tuning parameters * f16_f8_fp16 * cut test time * add f16_f8_fp16 * add f16_f8_f16 * testing instances for skinny cases * format * clean * add fp16_f8_fp16 * clang-format * add grouped gemm instalces * fixed profile grouped_gemm * clean * clean * clean * clean * clean * add missing instance func * fixed inferface --------- Co-authored-by: Jing Zhang <jizha@amd.com> Co-authored-by: root <root@sh5-1e707-rc06-38.mkm.dcgpu>	2023-11-07 09:09:58 -06:00
Illia Silin	f46a6ffad8	Fix the fp8 gemm for large tensors on MI300. (#1011 ) * Fix the fp8 conversion * Try clipping value before conversion * Fix return * Simplify with a const * reduce the gemm input tensor values to reduce round-off error * replace if-else with lambda * fix syntax --------- Co-authored-by: Rostyslav Geyyer <rosty.geyyer@amd.com>	2023-10-27 21:10:47 -07:00
Rostyslav Geyyer	1fd27d520f	Fix bf8 conversion issues (#1003 ) * Fix the conversion * Add bf8 functionality * Enable example on MI200 as well	2023-10-20 08:00:45 -05:00
Illia Silin	f7331c603b	Fix the DL kernel issues on Navi3x. (#998 ) * apply the patch for dl kernels on gfx11 * build DL kernels on navi32 CI	2023-10-19 09:34:39 -07:00
Bartłomiej Kocot	82f3a835d5	Extend available elementwise operations with conv examples (#995 ) * Extend available elementwise operations with conv examples * Fixes * Remove not needed convert * Update CMakeFile and dir name	2023-10-19 17:23:19 +02:00
zjing14	bf435140dc	Clean DTYPES conditions in CMake (#974 ) * Add a condition to build fp8 instances * simplified buffer_load/store * add bfp8/fp8 * fixed * remove all f8/bf8 condition include folder * fixed cmake conditions * fixed DTYPES=fp16/bfp16 * fix * fixed buffer_load * fixed buffer_store * fix * clean example cmake files * fixed ci * fixed cit --------- Co-authored-by: Rostyslav Geyyer <rosty.geyyer@amd.com> Co-authored-by: Jing Zhang <jizha@amd.com>	2023-10-18 11:14:14 -05:00
zjing14	39430bfdeb	workaround with float (#992 ) Co-authored-by: Jing Zhang <jizha@amd.com>	2023-10-16 15:42:59 -07:00
zjing14	2ce9b56c64	add vector_type support into thread_copy_v3r1 (#969 ) * add vector_type support into thread_copy_v3r1 * remove unncessary type_convert * fixed datatype * fixed dataType * changed API with is_packx_invocable * changed example * add missing cmake file * fixed ci * fixed cmake --------- Co-authored-by: Jing Zhang <jizha@amd.com>	2023-10-13 15:11:43 -05:00
zjing14	f3b02ecfd2	simplified buffer_load/store (#971 ) * simplified buffer_load/store * add bfp8/fp8 * fixed * fixed buffer_load * fixed buffer_store --------- Co-authored-by: Jing Zhang <jizha@amd.com>	2023-10-11 20:29:01 -05:00
zjing14	c99323be6e	Revert "Grouped Gemm with looping over the tiles. (#788 )" (#982 ) This reverts commit `a4f72a314a`.	2023-10-11 14:27:29 -05:00
Adam Osewski	a4f72a314a	Grouped Gemm with looping over the tiles. (#788 ) * Introduce LocalBlockToCTileMap. * Change the signature of CalculateBottomIndex() function which now does not accept any argument. The B2C map which is already passed as an argument to the kernel Run function is calculating block's local id already outside at kernel entry point __global__ function. The LocalB2C map stores as members local block ID. * Use LocalBlockToCTile map in device ops. * First draft of tile loop work distribution. * Fix typo. * Simplify kernel arguments. Calculate descriptors & B2C maps on the device. * Use looping kernel. * Fix B2C constructor. * Fix Navi21 errors. * Calculate tile start/end in device kernel. * Change Run API to accept user provided workspace buffer. * Add new line at EOF. * Move Gemm KernelArguments to device op interface. * Remove unused code. * Update API. * Launch grid size which is min of occupancy vs tile count * Get back to use constant memory for gemm descriptors. * Remove unused code. * Add default virtual method implementation. * Update comments to conform with doxygen style. * Fix doc style and unused parameters. * Add thread cluster lengths to kernel name. * Remove old splitk impl and replace it with tile looping one. * Modify instances. * set KPerBlock to 64 * maximize wherever possible vector load size. * Fix instances cluster lengths. * Change comment style. * Use 128b store where possible in instances. * Update test cases, since KPerBlock has doubled. * Update output stream operator for Sequence. * Add pipeline version to GroupedGEMM device op type string. * Fix pipeline version type logging. * Fix input tensors type after merge. * Fix compiler error. * Fix output stream operator for Pipeline version. * Store using 128b. * Set of instances with kpb 32/64 * Limit number of instances * Remove commented out instances. * Fix function name. * Limit the number of instances. Add pipline version to the regular instances * Change thr cluster layout for reading B tensor. * disabled failed instances --------- Co-authored-by: Adam Osewski <aosewski@amd.com> Co-authored-by: zjing14 <zhangjing14@gmail.com> Co-authored-by: Jing Zhang <jizha@amd.com>	2023-10-10 22:21:15 -05:00
zjing14	ac9595a9f1	Fixed f8_gemm NaN (#975 ) * workaround nan problem by changing output to fp16 * enable f8/bf8 gemm tests on MI200 * workaround f16 to f8 conversion --------- Co-authored-by: Jing Zhang <jizha@amd.com>	2023-10-10 10:30:26 -05:00
Rostyslav Geyyer	42facfc6b7	Add conv bwd weight fp16 comp bf8 fp8 op, instances and example (#945 ) * Add f8 bf8 gemm example * Add element-wise ops * Add intrinsics * Update reference calculation * Add an additional type option for xdlops gemm * Fix build process * Add bf8 to buffer addressing * Update blockwise op, split typeA and typeB * Update for compatibility * Uppdate naming to f8->fp8 * Update naming * Format * Update naming (#937) * Add a client example * Add computetypes to device and gridwise ops * Add instances, update instance factory * Format * Fix a flag * Add ckProfiler mode * Fix typos * Add an example * Add bf8 generator * add bf8 mfma; fixed type_convert for bf8 * move verfication ahead of timing * Update reference calculation * Fix reference * Narrow down float init range * Fix bf8 bf8 mfma * Add bf8 @ fp8 mfma * Update example * Update instances * Update profiler api * Update for compatibility * Format * Remove extra example * Clean up * workaround convert --------- Co-authored-by: Jing Zhang <jizha@amd.com>	2023-10-04 08:19:08 -05:00
Rostyslav Geyyer	bd09b5c538	Add fp8 @ bf8 gemm support and example (#933 ) * Add f8 bf8 gemm example * Add element-wise ops * Add intrinsics * Update reference calculation * Add an additional type option for xdlops gemm * Fix build process * Add bf8 to buffer addressing * Update blockwise op, split typeA and typeB * Update for compatibility * Uppdate naming to f8->fp8 * Update naming * Format	2023-10-02 16:39:03 -05:00
Bartlomiej Wroblewski	f4af5aed8b	Handle type conversions to a const datatype (#944 ) * Handle type conversions to a const datatype * Review: Handle X being const data type as well * Review: Remove typo	2023-09-27 15:02:42 -05:00
Bartłomiej Kocot	e2243a4d1e	Add column to image kernel (#930 ) * Add column to image kernel * Minor fixes for dtypes and client examples * Disable tests for disabled dtypes * Disable add instances functions for disabled data types * Minor stylistic fixes * Revert "Disable add instances functions for disabled data types" This reverts commit `728b869563`. * Instances reduction * Add comments in device_column_to_image_impl * Update changelog and Copyrights * Improve changelog	2023-09-27 17:19:06 +02:00
zjing14	11676c7e49	Add multiple A/B support (#906 ) * add gridwise_multi_abd * move element_op into RunRead * merge element_wise op with data read * add multiABD example * allow packed elementwise_op * changed example * clean * clean * add is_detected * fix * minor fix * add scaleAdd_vec4 example --------- Co-authored-by: Jing Zhang <jizha@amd.com>	2023-09-26 21:16:23 -05:00
Rostyslav Geyyer	f17af2e9ed	Add native conversions fp8<->fp32 (#908 ) * Add native conversions * Add bf8 conversions	2023-09-17 20:56:27 -05:00
Bartłomiej Kocot	475188ca2e	Add grouped conv bwd weight dl instances and new layout (#897 ) * Add grouped conv bwd weight dl instances and new layout * Add M and N padding * Remove todo comment * Enable grouped conv fwd dl k,c=1 generic instance * Comment fixes	2023-09-13 10:14:31 -05:00
Rostyslav Geyyer	62d4af7449	Refactor f8_t, add bf8_t (#792 ) * Refactor f8_t to add bf8_t * Add check_err impl for f8_t * Update fp8 test * Format * Revert the fix * Update vector_type implementation * Add bf8 test * Add bf8, use BitInt types * Add bf8 conversion methods * Update type_convert for fp8/bf8 * Add check_err fp8/bf8 support * Add subnorm fp8 tests * Add subnorm bf8 tests * Fix conversion * Add bf8 cmake bindings * Add macros to enable build with disabled fp8/bf8 * Remove is_native method * Update flag combination for mixed precision instances * Add more flag checks * Add another flag to a client example * Add type traits, decouple f8/bf8 casting * Clean up * Decouple fp8 and bf8 flags * Remove more redundant flags * Remove leftover comments	2023-09-12 17:04:27 -05:00
Bartlomiej Wroblewski	37a8c1f756	Redesign the DPP8 GEMM kernel to use warp-wise component (#863 ) * Redesign the DPP8 GEMM kernel to use warp-wise component * Review: Improve error messages * Review: Remove unnecessary empty lines * Review: Fix M, N per thread names * Review: Rename mfma_input_type to dpp_input_type * Review: Fix tensor adaptor; remove unnecessary element * Review: Remove calls to dpp_gemm's MakeCDescriptor * Review: Add blockwise doc, change function names to include dimension names * Review: Remove duplicated code; Move Block2CtileMap alias to the top of the file * Review: Add __restrict__ keywords * Review: Use MatrixPadder for padding A, B, C matrices * Review: Remove hardcoded datatypes * Review: Change names from FloatX to XDataType * Review: Introduce AK0 and BK0 instead of a single K0 * Review: Remove construction of dpp_datatypes object * Review: Rename DppInstrRunner to DppLanegroupGemm	2023-09-06 11:44:09 -05:00
rocking	866377de18	MaxPool & AvgPool bwd instances, test, ckProfiler, client example (#861 ) * Add maxpool instances * Rename index pool to max pool. * Add maxpool bwd bf16 instances * Add avg pool bwd instances * Rename avgpool and maxpool to avg_pool3d and max_pool * Add bf16 pool fwd instances * Add max pool bwd to ckProfiler * Add avg pool3d bwd to ckProfiler * Add avg pool bwd test * Fix bug of reference pool fwd (dilation) * Fix bug of max pool bwd (dilation and initZero) * Support bf16 compute data type * Force compute type be f32. Because atomicAdd only support f32 * Add max pool bwd test * Rename folder * Rename pool * Add max pool bwd client example * Add avg pool bwd client example * Add missing workspace * clang format * Rename macro * remove useless header * remove useless layout	2023-08-31 21:01:50 +08:00
Bartlomiej Wroblewski	32fe996da0	Fix datatype in inner_product when V_DOT2 is disabled (#849 )	2023-08-17 10:54:11 -05:00
Bartlomiej Wroblewski	d4c84256f7	Implement DPP8 based GEMM for Navi21 (#826 )	2023-08-14 15:46:27 -05:00
Bartłomiej Kocot	7761e5232c	Add s_nops after v_dot to avoid hazard (#808 ) * Add s_nops after v_dot to avoid hazard * Fix builtin for inner_produxt fp16 * Skip inline version to builtin * Add comments regarding isa * Fix comment regarding s_nop	2023-07-27 13:29:44 -05:00
carlushuang	e7dca79d27	initial stream-k implementation with example (#699 ) * initial stream-k implementation with example * fix unexpected change in err * improve a little bit performance by reorganize pipeline. * improve perf a little bit by swizzle block idx * add profiler * update example * fix spelling * shrink karg for streamk * support dynamic buffer using memory coherence glc_slc bit from template * control memory coherence while construct dynamic buffer * update reduction for streamk(not ready yet) * Add template parameter to make_dynamic_buffer to support amd_buffer coherence setting * fix build issue * fix several bug * now result is correct, everything works (but has scratch) * remove scratch by manually reset coordinate * update device code * fix a bug in final reduce * fix something in example * update async memset * fix enum as camel case * modify coherence enum name * clean code and use atomic streamk by default * remove unused var * throw exception if have empty pointer * fix format * fix CI warning * fix type in init * modify CI error * filter out on gfx10+ * restore changed example code --------- Co-authored-by: Qianfeng Zhang <Qianfeng.Zhang@amd.com>	2023-07-26 14:18:15 -05:00
Rostyslav Geyyer	f82bd59389	Remove type_convert bf16 to int32 and back (#802 )	2023-07-18 09:44:51 -05:00
Qianfeng	8f5cafaf04	Batchnorm splitk single kernel (#771 ) * Use dim 0 as faster dim for writing mean/var/count workspace in batchnorm multiblock method [performance] * Add CountDataType as template parameter in blockwise_welford * Add utility/get_shift.hpp * Add BatchNorm multiblock single-kernel implementation * Add smem inline assembly based implementation of gms_init/gms_barrier/gms_reset for gfx90a * Renaming in device_batchnorm_forward_impl.hpp * Tiny fix in the batchnorm_fwd profiler * Revert "Add smem inline assembly based implementation of gms_init/gms_barrier/gms_reset for gfx90a" This reverts commit `d16d00919c`. * Use the old two-kernel batchnorm multiblock method for gfx1030 * Use the old two-kernel batchnorm multiblock method for gfx908 * use the single-kernel batchnorm multiblock method only for gfx90a * Remove get_wave_id() from utility/get_id.hpp since it is not used * Set true for testing running mean/variance and saving mean/invvariance in the examples * Fix to copy-right words * Remove un-needed including in utility/get_id.hpp * Add comments to workgroup_synchronization.hpp * Remove un-used codes in gridwise_multiblock_batchnorm_forward.hpp * Renaming in the kernels * Remove un-used kernel file	2023-07-06 10:58:55 -05:00
Rostyslav Geyyer	61dc9aa932	Add the missing archs (#785 )	2023-07-05 18:29:56 -05:00
Rostyslav Geyyer	1cf5003179	Add fp8 GEMM and an example for it (#767 ) * Add fp8 xdl gemm * Add example * Use int8 intrinsics for buffer load/store * Format * Update cmakelists	2023-07-04 20:38:49 -06:00
Rostyslav Geyyer	f0c620c42e	FP8 enablement - add a pseudorandom number generator, add conversion methods (#708 ) * Add basic fp8 definitions and prn-generator * Format * Add fp8<->fp32 type_convert * Format * Split type_convert and cast_to/from_f8 * Format * Minor fix * Minor fix * Move fp8 utils to a separate header * Add elementwise ops * Add fp8_convert_sr * Format * Add element op * Eliminate magic numbers * Split f8_convert_sr in host and device * Format * Add some constexpr * Add a datatype test * Format * Another format * Add fp8<->fp16 tests * Update type_converts * Format * Add fp16 casting functions * Format * Use seed as a runtime arg * Use element location for PRNG * Format * Add fp8<->fp16 to PassThrough element op * Clean up * Merge host and device implementations * Add comments on rounding modes * Remove leftover code * Put type_converts into a separate header * Put random number gen to a separate header * Rearrange f8_utils' namespaces * Refactor type_convert.hpp * Move f8_t definition	2023-06-19 11:20:35 -05:00
Illia Silin	027e46ee82	Enable gfx941 and gfx942 architectures. (#752 ) * enable gfx941/942 targets * fix clang format * fix the cmake logic for multiple targets * fix cmake syntax for looping over targets * add gfx941/942 support for gemm_xdl instances	2023-06-15 08:20:59 -07:00
Po Yen Chen	7c24654c24	Fix incomplete object size (=4n + 3) support of amd_wave_read_first_lane() (#738 ) * Fix wrong pointer type * Rename type trait get_unsigned_int<> to get_carrier<> * Add 3-bytes carrier type * Add missing __device__ specifier * Rename template non-type parameter * Leave the rest byte uninitialized * Avoid invoking (host) STL algorithms * Remove unnecessary 'inline' specifier * Extract common logic out as helper method * Hide dummy member function * Add missing __device__ specifier	2023-06-12 08:36:40 -05:00
carlushuang	016ebaa7f3	support dynamic buffer using memory coherence glc_slc bit from template (#725 )	2023-06-08 07:40:29 -05:00
Illia Silin	b94fd0b227	update copyright headers (#726 )	2023-05-31 18:46:57 -05:00
Po Yen Chen	582e31e88d	Add class type support for __builtin_amdgcn_readfirstlane() (#711 ) * Add overloaded version of __builtin_amdgcn_readfirstlane() * Remove 'static' specifiers * Remove more 'static' specifier * Replace unsigne char by std::byte * Add 'const' specifier to never changing variable * Add 'inline' specifier to funcion definition * Fix wrong boundar calculation logic * Rename type trait * Remove std:: qualifier from standard types * Replace 'size_t' by 'unsigned' * Use type alias to hint usage * Replace static_for<> by ordinary 'for' loop * Rename readfirstlane() to amd_wave_read_first_lane() * Rename file readfirstlance.hpp as amd_wave_read_first_lane.hpp * Reorder statements	2023-05-31 10:25:25 -05:00
Illia Silin	ac9e01e2cc	Clean-up the headers (#713 ) * fix headers for gpu instances * remove unused headers --------- Co-authored-by: zjing14 <zhangjing14@gmail.com>	2023-05-24 08:11:25 -07:00
Rostyslav Geyyer	b076a02ad2	Optimize bf16 conversion (#664 ) * Add TypeConvert class and start refactoring * Refactor TypeConvert as a struct * Get back to template functions type_convert * Add a type_convert_bf16_rtn, set rtz as default * Clean up * Add UnaryConvertPrecision struct for high-precision workloads * Format * Update type_convert to UnaryConvert on threadwise level * Update UnaryConvertPrecision * Format * Fix chmod * Add a flag to pick converion method * Format * Remove the added flag * Merge elementwise op with type conversion * Move type_convert to elemwise op, update the op * Update type_convert_precision -> bf16_convert_rtn * Clean up * Update comments * Update the CK_WORKAROUND_DENORM_FIX flag handling * Update the unneeded op to work but warn user * Remove the message * Use a PassThrough instead of ConvertBF16RTN to calcaulate reference * Format * Add missing include	2023-05-04 10:25:47 -05:00
Illia Silin	4feebedd41	Syncing up from internal repo to enable MI300. (#690 ) * enable gfx940 * switch between intrinsic mfma routines on mi100/200 and mi300 * fix mfma_int8 on MI300 * disable 2 int8 examples on MI300 * Update cmake-ck-dev.sh * restore gitignore file * modify Jenkinsfile to the internal repo --------- Co-authored-by: Jing Zhang <jizha@amd.com> Co-authored-by: zjing14 <zhangjing14@gmail.com>	2023-04-28 18:22:59 -05:00
rocking5566	ed3a2e5226	Groupnorm + swish external api (#668 ) * Rename to proper naming * Add example of groupnorm + swish * Extract duplicate code in example * Add groupnorm + swish instances * Ractor instance generation, split into multiple cpp file * Add external api and client example * Refine profiler message * Use ck math version of exp * Refine problem size in example * Add host version of exp	2023-04-10 08:02:17 -05:00
Rostyslav Geyyer	dbd8f94bef	Add a denorm test fix (#603 ) * Add type_convert implementations for bf16 * Add the fix for conv_fwd * Add the fix for conv_bwd_data * Add the fix for conv_bwd_weight * Format * Format * Another format * Add a macro to use workaround on MI200 only * Format --------- Co-authored-by: Rosty Geyyer <rosty.geyyer@amd.com> Co-authored-by: zjing14 <zhangjing14@gmail.com>	2023-03-29 15:05:32 -05:00

1 2 3

108 Commits