composable_kernel

mirror of https://github.com/ROCm/composable_kernel.git synced 2026-06-07 16:26:10 +00:00

Files

John Shumway 37dff024c1 [CK_BUILDER] Add compile-time reflection for a convolution instance (#3065 )

* [CK_BILDER] Add compile-time reflection for a convolution instance

Introduce InstanceTraits template metaprogramming framework to enable runtime introspection of device kernel template parameters without requiring implementation knowledge. This reflection system extracts configuration details (block sizes, data types, layouts, tuning parameters) directly from kernel specializations through template
pattern matching. In particular, the GetInstanceString method returns a string that uniquely idenitfies the kernel, by explicitly serializing all template paramter values.

This provides critical functionality for MIOpen integration, since the existing GetTypeString method is ambiguous, and only captures some of the template paramters.

The implementation uses a two-level design: a primary InstanceTraits template declaration in instance_traits.hpp serves as the interface, while kernel-specific specializations (e.g., for DeviceGroupedConvFwdMultipleABD_Xdl_CShuffle_V3) provide the actual extraction logic. This separation allows the reflection system to scale to additional kernel types without modifying the core interface.

Key architectural decisions:

- Forward-declare device kernels in instance_traits.hpp to avoid  circular dependencies, since device implementation headers will  include the reflection headers

- Use compile-time constants and type aliases to expose kernel  parameters, enabling zero-overhead introspection

- Provide a templated instance_string() function that generates human-readable  kernel configuration strings by serializing all template parameters  in order, useful for debugging and kernel identification

- Guard reflection integration with preprocessor definition CK_EXPERIMENTAL_BUILDER to keep  it opt-in until the API stabilizes

- Add GetInstanceString() virtual method to BaseOperator, allowing  runtime polymorphic access to compile-time kernel information

This infrastructure also enables upcoming higher-level semantic reflection abstractions (like ConvTraits) to query kernel configurations programmatically.

Includes unit tests validating both the trait extraction accuracy and the string generation format.

2025-10-21 21:10:19 -07:00

block

Wave Tile Transfer supporting global load with transpose (#3027 )

2025-10-16 11:33:56 -07:00

device

[CK_BUILDER] Add compile-time reflection for a convolution instance (#3065 )

2025-10-21 21:10:19 -07:00

element

Wmma support for multiple Ds based GEMMs (#2613 )

2025-09-05 16:31:08 +02:00

grid

Gridwise gemm conv v3 force padded layout on gfx950 (#2961 )