Adding Instances and Examples for FP8-based Scaled Convolution and AMAX Reduction. (#1473)

* Enable CMakePresets build * Verify Convolution, Scaling and ReLU algorithms. * Add tensor element-wise scale and type cast operation. * Reduction implemented but does not work. * Exploration of Reduction functionality. * Completed example for Convolution scaled with ReLu activation and AMAX reduction. * WIP: Add required instances for convolution. * WIP: Create client example. Implement convolution stage. * Add elementwise instances. * Add elementwise scale + convert example. * Add reduction instances. * WIP: Client example for AMAX reduction. * WIP: Add instances for multistage reduction. * WIP: Implementation of multistage reduction. * Refactoring. * Clean up. * Add CMakePresets.json * Guard off FP8 instances when the data type is not available. * Add example for Scaled FP8 Convolution with AMAX reduction. * Refactor CombConvScaleRelu instances. * Add CombConvScale instances. * Add client example for Scaled FP8 Convolution with AMAX reduction. * Cleanup. [ROCm/composable_kernel commit: c3515f277c]
2026-05-18 20:09:25 +00:00 · 2024-08-21 16:22:41 -06:00
parent 94954e9fe4
commit 10be209218
14 changed files with 389 additions and 87 deletions
--- a/include/ck/tensor_operation/gpu/element/combined_element_wise_operation.hpp
+++ b/include/ck/tensor_operation/gpu/element/combined_element_wise_operation.hpp
@@ -3,7 +3,6 @@

 #pragma once

-#include "ck/utility/data_type.hpp"
 #include "ck/tensor_operation/gpu/element/element_wise_operation.hpp"

 namespace ck {
@@ -107,6 +106,9 @@ struct TrinaryWithUnaryCombinedOp
    UnaryOp2 unary_op2_{};
 };

+using ScaleScalePass = UnaryCombinedOp<Scale, Scale, PassThrough>;
+using ScaleScaleRelu = UnaryCombinedOp<Scale, Scale, Relu>;
+
 } // namespace element_wise
 } // namespace tensor_operation
 } // namespace ck