mirror of https://github.com/ROCm/composable_kernel.git synced 2026-07-11 09:40:51 +00:00

Files

Thomas Ning f240ae3248 Enable Async Copy for MI355 (#2425 )

* add for async load builtin

* add async load api

* fix some compiling errors

* fix a compiling error

* fix some compiling errors

* add a pipeline which copies from v4

* add a new pipeline for async load

* fix some compiling errors

* add async load tests

* fix some issues in async load

* fix

* fix async inline assembly

* fix async inline assembly

* add ignore header file

* comment some not gfx950 codes

* comment some not gfx950 codes

* fix a error

* update async load apis

* fix lds descriptor

* fix a compiling error

* fix some compiling errors

* fix a descriptor issue

* update lds descriptor

* change async pipeline's tile distribution pattern from thread to warp

* fix clang format

* update async policy

* fix a CRTP issue

* fix a typo error

* change lds layout

* fix some sync issues

* improve codes

* delete the async test

* fix a commented format issue

* avoid compiling device functions when compile host

* make gemm run

* add the copy kernel support

* finish the feature

* Address comment

* add the support for buffer_builtin

* solved the merging problem

* Comment Addressed

---------

Co-authored-by: joye <joye@amd.com>
Co-authored-by: joyeamd <John.Ye@amd.com>

2025-07-07 10:08:49 -07:00

algorithm

[CK-TILE] File-level documentation for static encoding pattern (#2433 )

2025-07-04 02:26:18 -07:00

arch

Enable Async Copy for MI355 (#2425 )

2025-07-07 10:08:49 -07:00

container

[CK_TILE][CORE] enhance slice_tile api (#2430 )

2025-07-06 20:13:12 -07:00

numeric

[CK Tile] Int8 Support on CK Tile GEMM (#2267 )

2025-06-25 08:20:35 -07:00

tensor

Enable Async Copy for MI355 (#2425 )

2025-07-07 10:08:49 -07:00

utility

[CK_TILE] Tileloop persistent gemm - resubmit (#2299 )

2025-06-06 14:18:49 -07:00

config.hpp

default skip y point to r (#2457 )

2025-07-06 23:54:34 -07:00

README.md

introducing ck_tile! (#1216 )

2024-04-15 19:27:12 -05:00

README.md

ck_tile/core

ck_tile/core contains every basic functions and structures to create a GPU kernel using ck_tile. User should only include ck_tile/core.hpp this single header to use all the functionality. Everything is under ck_tile namespace. The coding style under this folder should be similar to std (snake_case for structure/function, Camel for template types...)

algorithm/
    coordinate transform and some other reusable algorithm
arch/
    contains some basic device building block like mma, buffer addressing, etc...
container/
    contains basic container data structure, array/sequence/tuple/...
numeric/
    data type, and data type related math
tensor/
    tensor descriptors and tile level API
utility/
    other utility function for both host/device