composable_kernel

mirror of https://github.com/ROCm/composable_kernel.git synced 2026-05-11 17:00:18 +00:00

Author	SHA1	Message	Date
Bartłomiej Kocot	31bf253aeb	Add dynamic elementwise op (#1426 ) * Add dynamic elementwise op Co-authored-by: ThruptiRajLakshmanaGowda <thruptiraj.lakshmanagowda@amd.com> * CI issues fix * Custom parameter value for dynamic functions - Comments addressed --------- Co-authored-by: ThruptiRajLakshmanaGowda <thruptiraj.lakshmanagowda@amd.com> Co-authored-by: ThruptiRajLakshmanaGowda <tlakshma@amd.com>	2024-10-26 15:22:37 +02:00
Po Yen Chen	54f0e6f4bb	[CK_TILE] More fmha splitkv optimizations (#1588 ) * Use pre-defined constants for readability * Use vector write for o_acc tensor * Remove no-longer used policy method * Deprecate no-longer used policy/pipeline * Specify gemm0/gemm1 block warps separately in codegen * Fix wrong ps_idx creation logic * Add single-warp block gemm * Supoprt single-warp gemm0 * Make MakeCBlockTile() as static method * Use MakeCBlockTile() to get underlying tile distribution * Use kNumGemm1Warps to compute # threads for gemm1 * Put normal case in the if clause * Refine fmha splitkv block mapping * Refine & fix the lse_acc/o_acc layout * Fix wrong LDS size for K tile * Use kK0=64 for hdim=128,256 fmha splitkv kernels * Use kK1=64 for hdim=32,64,128 fmha splitkv kernels * Undo kK0/kK1 changes * Use more reasonable GetAlignmentV() computation * Using store_tile() in fmha splitkv kernel epilogue	2024-10-26 18:35:45 +08:00
valarLip	37f7afed1e	add int8 gemm multiply multiply a8w8 (#1591 ) * add int8 gemm multiply multiply a8w8 * uncomment * clang-format-12 * Add example_gemm_multiply_multiply_xdl_int8 * Remove shell scripts * update preprocess number for mi308; bring back printout in ckprofiler * format --------- Co-authored-by: chenjun <junchen2@amd.com> Co-authored-by: Haocong WANG <haocwang@amd.com> Co-authored-by: carlushuang <carlus.huang@amd.com>	2024-10-26 16:39:34 +08:00
Max Podkorytov	eda5938386	add parsing grouped conv fwd instances	2024-10-25 08:25:53 -07:00
Rostyslav Geyyer	7d576f1748	Update GPU verification (#1596 ) * Update inits * Update static_cast to type_convert * Add verification option selection	2024-10-25 08:13:46 -07:00
aledudek	9385caa306	Generic threshold calculation (#1546 ) * Calculate generic relative threshold pool3dfwd * Calculate absolute error threshold pool3d fwd * Generic threshold calculation take max input for relative error pool3dfwd * Remove max possible value for error calculation at runtime * Remove debug print in pool3dfwd * Pool3d fwd adjusted types in generic threshold calculation * Generic threshold calculation take into account number of accumulations and accdatatype * Generic threshold fix final error formula * Generic threshold calculation - num of accs fix * Generic threshold calculation - adjust absolute error * Generic threshold calculation - OutDataType in absolute error	2024-10-25 12:46:24 +02:00
dummycoderfe	9183ce69ca	hot_fix epsilon pos (#1597 ) Co-authored-by: dummycoderfe <noplydummmycoder@163.com>	2024-10-25 11:17:45 +08:00
Illia Silin	8e22e1ae31	fix the logic of enabling XDL and WMMA instances (#1595 )	2024-10-23 15:55:39 -07:00
Bartłomiej Kocot	cedccd59c9	[POST MERGE PR] Enable grouped conv bwd wei bf16 NGCHW (#1594 )	2024-10-23 12:02:33 +02:00
Jatin Chaudhary	4d5248e2d1	Explicit cast values to half (#1593 ) Co-authored-by: Illia Silin <98187287+illsilin@users.noreply.github.com>	2024-10-22 11:17:32 -07:00
Bartłomiej Kocot	82fc53835a	Enable grouped conv bwd wei bf16 NGCHW (#1589 ) * Enable grouped conv bwd wei bf16 NGCHW * fixes * fixes * Fixes * fixes * fixes * Fixes	2024-10-22 16:18:28 +02:00
ltqin	0394f8a713	update layernorm (#1570 ) * port layernorm * change warp_welford.hpp * Update warpshuffle * 1. Add save mean and save std back 2. Move construction of tensor_view and tile_window to operator() * refine welford max count calculation * unify layernorm api * Rename file * Remove save mean and inv std * Revert "refine welford max count calculation" This reverts commit `022365802b`. * Fix order of parameter * refine welford max count calculation again * Remove fp32 instances * Fix bug of padding * refactor api * Support bf16 * Extract common function * Refine arg of operator() * Add kMThreadPerBlock to template parameter * clang format * Refine variable name * Refine file name * remove redundant line * refactor layernorm2d pipeline and add block-per-block utility * fix name * rename more * add more block-per-tile instance * remove duplicated define * update instance for 2048, 1024 case * support up to 2048 now * opt loading * add n1536 * Add two pass pipeline * format * Fix incorrect type * parallel compilation * Use smaller N * fix 2p pass * Support Repeat_M in distribution * Refine nameing * Add reduce example --------- Co-authored-by: letaoqin <letaoqin@amd.com> Co-authored-by: aska-0096 <haocwang@amd.com> Co-authored-by: rocking <ChunYu.Lai@amd.com> Co-authored-by: carlushuang <carlus.huang@amd.com>	2024-10-22 09:26:18 +08:00
Rostyslav Geyyer	3f710930f6	Update default stride (#1576 ) * Update default stride value to -1 * Fix format * Revert "Fix format" This reverts commit `ae0c3649ec`. --------- Co-authored-by: Harisankar Sadasivan <135730918+hsadasiv@users.noreply.github.com>	2024-10-21 08:45:22 -07:00
spolifroni-amd	794f2d64a8	added link to documentation (#1578 )	2024-10-21 08:35:57 -07:00
dependabot[bot]	d0565e33d6	Bump rocm-docs-core from 1.8.2 to 1.8.3 in /docs/sphinx (#1587 ) Bumps [rocm-docs-core](https://github.com/ROCm/rocm-docs-core) from 1.8.2 to 1.8.3. - [Release notes](https://github.com/ROCm/rocm-docs-core/releases) - [Changelog](https://github.com/ROCm/rocm-docs-core/blob/develop/CHANGELOG.md) - [Commits](https://github.com/ROCm/rocm-docs-core/compare/v1.8.2...v1.8.3) --- updated-dependencies: - dependency-name: rocm-docs-core dependency-type: direct:production update-type: version-update:semver-patch ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>	2024-10-21 08:34:53 -07:00
Thomas Ning	560917b161	Ck profiler instance support (#1575 ) * The draft on ckProfiler instance add * support the ck profiler instance with same data types * add a small feature on the M and N variable switch. * Partially solve the incorrect result problem * fix based on ci cd	2024-10-21 22:47:48 +08:00
Po Yen Chen	95e722a3b3	[CK_TILE] Optimize fmha splitkv & splitkv combine kernels (#1577 ) * Use smaller width for lse_accum dist tensor * Update pipeline comment * Fix wrong distribution for lse_accum * Remove duplicate dim in lse_accum dist encoding * Decide fmha splitkv combine kernel kBlockSize by kM0 * Remove assumption of MPerThread=1 * Add log<4> & log<8> specialization * Enlarge occupancy array * Fix vector size for small tile * Add support for kMaxSplits=8 * Re-format gemm.hpp * Use 16x16x16 warp gemm for fwd_splitkv * Centralize policy code changes * Leave fp8/bf8 tile settings unchanged	2024-10-21 10:52:11 +08:00
Haocong WANG	a285d6f9b5	disable bad instance detected on MI308CPX (#1584 )	2024-10-18 08:46:11 -07:00
Illia Silin	88e6fa7fdb	add the lsr-drop-solution=1 compiler flag (#1582 )	2024-10-18 08:25:54 -07:00
Qianfeng	14c3cfb1c6	[CK_TILE] Improve headdim96 performance for fmha-bwd (#1573 ) * Add kQKHeaddimForGemmN and kVHeaddimForGemmN in order to support headdim 96 * Remove the using of MakeKRegBlockDescriptor and MakeVRegBlockDescriptor * Fix in bwd_piple_default_policy * Remove kQKHeaddim and rename kQKHeaddimForGemmN to kQKHeaddim in the bwd kernel and pipelines * Replace kVHeaddimForGemmN by kVHeaddim and kDoDvHeaddim * Update to hd96 tile settings * Add smoke test scripts for fmha-bwd hd96 * Revert "Add smoke test scripts for fmha-bwd hd96" This reverts commit `7ca7e1a93d`. * Remove hd96 tile settings in fmha_bwd codegen to save compiling * Fix lost code line in bwd_pipeline_default_policy * Merge kDoDvHeaddim/kPadHeadDimDoDv to kVHeaddim/kPadHeadDimV and remove TileFmhaBwdTraits * Rename KRegSliceBlockDescriptor/VRegSliceBlockDescriptor to KRegBlockDescriptor/VRegBlockDescriptor * tiny adjustments --------- Co-authored-by: Po Yen Chen <PoYen.Chen@amd.com> Co-authored-by: danyao12 <Dan.Yao@amd.com>	2024-10-16 18:14:32 +08:00
Paul Fultz II	10158b0ffd	Build codegen as standalone (#1556 ) * Build codegen as standalone * Add exception for device tests * Use local filesystem header * add a codegen test CI stage and daily build --------- Co-authored-by: illsilin <Illia.Silin@amd.com> Co-authored-by: Illia Silin <98187287+illsilin@users.noreply.github.com>	2024-10-15 13:20:42 -07:00
Bartłomiej Kocot	d02a92cc0d	[CK_TILE] Add block universal gemm pipeline policy (#1557 ) * [CK_TILE] Add block universal gemm pipeline policy * Fixes * fixes2 * Fixes3 * fixeS	2024-10-15 13:53:41 +02:00
Po Yen Chen	9868fd0245	Apply ROCm 6.2 WA to ROCm 6.3 and later (#1563 )	2024-10-15 18:02:41 +08:00
Rostyslav Geyyer	4cf70b36c1	Add custom type vector support (#1333 ) * Add non_native_vector_type * Add a test * Add non-native vector type * Fix CTOR * Fix non-native vector type of 1 * Fix CTORs * Use vector_type to cover non-native implementation as well * Update the test * Format * Format * Fix copyright years * Remove BoolVecT so far * Add AsType test cases * Update assert error message * Remove redundant type * Update naming * Add complex half type with tests * Add tests for vector reshaping * Add missing alignas * Update test/data_type/test_custom_type.cpp Co-authored-by: Adam Osewski <19374865+aosewski@users.noreply.github.com> * Compare custom types to built-in types * Add default constructor test * Add an alignment test --------- Co-authored-by: Illia Silin <98187287+illsilin@users.noreply.github.com> Co-authored-by: Adam Osewski <19374865+aosewski@users.noreply.github.com> Co-authored-by: Po Yen Chen <PoYen.Chen@amd.com>	2024-10-14 11:56:45 -05:00
Bartłomiej Kocot	f21cda2536	Add transpose scale amax example (#1547 ) * Add transpose scale amax example * fixes * Tune reduce instance	2024-10-14 17:39:38 +02:00
Thomas Ning	35c1777d59	decouple the calling from gemm_pipeline (#1571 ) * decouple the calling from gemm_pipeline * clang format	2024-10-14 13:59:26 +08:00
Adam Osewski	29d384d0b2	Implement GetWorkSpaceSize from BaseOperator. (#1564 )	2024-10-12 14:05:11 +08:00
Illia Silin	11444e4cf2	[CI] remove the --rm docker container flags (#1568 )	2024-10-11 14:29:46 -07:00
Illia Silin	f46a9eee9d	only build tests and examples if user sets GPU_TARGETS (#1565 )	2024-10-10 15:31:56 -07:00
spolifroni-amd	14c52befda	removed API usage header (#1566 )	2024-10-10 13:57:23 -07:00
Rostyslav Geyyer	d18fc0797f	Fix default stride value (#1559 )	2024-10-10 07:37:09 -07:00
Thomas Ning	6f27bc9872	Ck tile gemm cshuffle & CK Tile GEMM restructure (#1535 ) * ake the cshuffle compilable * modify Mhe reference on gpu and cpu. Correaccess of cshuffle * fix the cpu reference code * Complete the in tile shuffle logic * restructure the kernel template input * change the naming pattern of ck_tile gemm pipeline * Re-format files using remod.py * Solve the fmha conflict with gemm * Comment Addressed from Carlus --------- Co-authored-by: Po Yen, Chen <PoYen.Chen@amd.com>	2024-10-10 18:02:22 +08:00
Illia Silin	2e1165c1a7	fix the target selection logic (#1561 )	2024-10-09 15:21:57 -07:00
Illia Silin	cfac9497e2	remove gfx12 targets from daily builds with rocm6.2 (#1560 )	2024-10-09 10:18:05 -07:00
Christopher Millette	ceaed8e097	Fixes small memory leak from missing hipEventDestroy (#1554 )	2024-10-09 09:41:35 +02:00
Rostyslav Geyyer	aa932445ea	Add a gpu gemm reference kernel (#1528 ) * Add a gpu gemm reference kernel * Switch to gpu reference in gemm examples * Remove redundant arguments * Update all related examples * Update more examples * Try less threads per block * Try even less threads per block * Add support for all matrix layouts * Increase block size * Clean up * Remove hardcoded strides * Clean up * Try a column-major case * Revert back to row-major * Run both CPU and GPU veriffication --------- Co-authored-by: Po Yen Chen <PoYen.Chen@amd.com>	2024-10-08 11:05:28 -05:00
Po Yen Chen	0c094daa7e	[CK_TILE] Update example README files & fix script compatibility issue (#1548 ) * Fix text alignment of ArgParser::print() * Update example README files * Clarify make-ck-dev.sh <arch> usage * Only keep some of the argument from '-?' output * Undo command line output changes in README * Only keep existing argument on doc and update description * Fix text alignment * Make cmake-ck-*.sh compatible with 'sh' command	2024-10-08 10:45:12 +08:00
Qianfeng	74d68e3b99	[CK_TILE] Simplify the codes in splitkv_combine pipeline (#1549 ) * Simplify the codes in splitkv_combine pipeline * Always set kPadSeqLenK=true for fmha splitkv kernels * Change in Oacc Alignment and TileDistribution to be more adaptable to tile sizes --------- Co-authored-by: Po Yen Chen <PoYen.Chen@amd.com>	2024-10-08 10:44:34 +08:00
Illia Silin	7733ae167b	add a CK_USE_CODEGEN build argument to enable codegen (#1552 ) * add a CK_USE_CODEGEN build argument to enable codegen * fix cmake codegen logic	2024-10-07 15:45:19 -07:00
Illia Silin	7d8ea5f08b	Fix build logic using GRU_ARCHS. (#1536 ) * update build logic with GPU_ARCHS * fix the GPU_ARCHS build for codegen * unset GPU_TARGETS when GPU_ARCHS are set	2024-10-07 08:18:23 -07:00
Bartłomiej Kocot	cc8f466a7e	[CK_TILE] Fix conv param multiple definition (#1550 ) Co-authored-by: Po Yen Chen <PoYen.Chen@amd.com>	2024-10-07 15:21:21 +02:00
rocking	0023f01ab0	[Ck tile] Support layernorm one pass (#1512 ) * Fix compile error * Add one pass pipeline * Extract creating tile_window to operator() * clang format * reduce duplicated code * do not hardcode * Support padding in layernorm --------- Co-authored-by: Po Yen Chen <PoYen.Chen@amd.com>	2024-10-07 14:25:53 +08:00
kylasa	c24fae2346	Adding seed and offset pointer support to the philox random number generator. (#1523 ) * Adding seed and offset pointer support to the philox random number generator. * Separating seed and offset pointer checks with different condition statements. * Changes include, adding support for device seed and offset pointers, union is used to store seed/offset values and device pointers to minimize device SGPRs. * Correcting a typo in the readme file * Re-format files using remod.py * Use STL type for API parameters * Use simpler struct design for drop_seed & drop_offset * Undo unnecessary changes * Sync kargs style for fmha_fwd.hpp/.cpp * Use templated union to reduce code * Use structured binding to make code more readable --------- Co-authored-by: Sudhir Kylasa <sukylasa@amd.com> Co-authored-by: Po Yen Chen <PoYen.Chen@amd.com>	2024-10-05 02:48:47 +08:00
arai713	b545de175a	Codegen build (#1526 ) * updating codegen build for MIOpen access: adding .cmake for codegen component (cherry picked from commit `652a7c0463`) * updating CMake (cherry picked from commit `a685822e36`)	2024-10-04 10:51:50 -07:00
Bartłomiej Kocot	6b54d2faf8	Fix grouped gemm check to avoid overflow (#1545 )	2024-10-04 17:32:43 +02:00
macurtis-amd	aeb7c91f48	Fix compilation errors generated by forthcoming Clang changes (#1544 ) Without this change, the following diagnostic is generated: a template argument list is expected after a name prefixed by the template keyword [-Wmissing-template-arg-list-after-template-kw] See C++17 spec [temp.names] p5.	2024-10-02 13:56:22 -07:00
BrianHarrisonAMD	294cb82314	Add generating mha static library for gfx90a (#1540 ) * Add generating mha static library for gfx90a * Update comment to reflect changes	2024-10-02 09:26:11 -07:00
Illia Silin	11b7a4db00	re-enable the FMHA performance monitoring (#1539 )	2024-10-01 13:17:55 -07:00
Illia Silin	8e4c3fb1bc	[CK_TILE] add missing vector header (#1537 ) * add missing vector header * Re-format header using remod.py --------- Co-authored-by: Po Yen, Chen <PoYen.Chen@amd.com>	2024-10-01 07:58:20 -07:00
Po Yen Chen	a1c07e8d91	[CK_TILE] Change output accum tensor layout of fmha fwd split-kv & combine kernels (#1527 ) * Use same layout for o_acc and o tensor * Use better param names in partitioner * Remove redundant kargs 'max_seqlen_q' * Use better param names in splitkv kernel * Add comment for additional kernel arguments * Sync empty loop early return logics between pipelines * Pass more arguments to cmake in scripts * Align backslashes * Fix wrong o_acc tensor view strides * Change o_acc layout if o_perm=0 * Handle whole row masked via attn_bias * Use use vector width = 1 for o_acc * Use more even split sizes	2024-10-01 22:13:52 +08:00

1 2 3 4 5 ...

1483 Commits