amd/blis - blis - Public git mirror

amd/blis

mirror of https://github.com/amd/blis.git synced 2026-07-02 21:27:52 +00:00

Author	SHA1	Message	Date
V, Varsha	8fd7060b2f	Matrix Add and Matrix Mul Post-op addition in F32 AVX512_256 kernels (#50 ) Added Matrix-mul and Matrix-add postops in FP32 AVX512_256 GEMV kernels - Matrix-add and Matrix-mul post ops in FP32 AVX512_256 GEMV m = 1 and n = 1 kernels has been added. Co-authored-by: VarshaV <varshav2@amd.com>	2025-06-17 16:17:13 +05:30
S, Hari Govind	e097346658	Implemented Multithreading Support and Optimization of DGEMV API (#10 ) - Implemented multithreading framework for the DGEMV API on Zen architectures. Architecture specific AOCL-dynamic logic determines the optimal number of threads for improved performance. - The condition check for the value of beta is optimized by utilizing masked operations. The mask value is set based on value of beta, and the masked operations are applied when the vector y is loaded or scaled with beta. AMD-Internal: [CPUPL-6746]	2025-06-17 12:39:48 +05:30
V, Varsha	875375a362	Bug Fixes in FP32 Kernels: (#41 ) * Bug Fixes in FP32 Kernels: - The current implementation lets m=1 tiny cases inside LPGEMV_TINY loop, but the m=1 GEMV kernel call doesn't have the call to GEMV_M_ONE kernels. Added the m=1 path in LPGEMV_TINY loop by handling the pack A/Pack B/reorder B conditions. - Added BF16 support for BIAS, Matrix-Add and Matrix-Mul for AVX512 F32 main and GEMV kernels - Added BF16 Matrix-Add and Matrix-Mul support for AVX512_256 F32 kernels. - Modified the condition check in FP32 Zero point in AVX512 kernels, and fixed few bugs in Col-major Zero point evaluation. AMD Internal: [ CPUPL - 6748 ] * Bug Fixes in FP32 Kernels: - The current implementation lets m=1 tiny cases inside LPGEMV_TINY loop, but doesn't have the call to GEMV_M_ONE kernels. Added the m=1 path in LPGEMV_TINY loop by handling the pack A/Pack B/reorder B conditions. - Added BF16 support for BIAS, Matrix-Add and Matrix-Mul for AVX512 F32 main and GEMV kernels. - Added BF16 Downscale, BIAS, Matrix-Add and Matrix-Mul support in AVX2 GEMV_N and AVX512_256 GEMV kernels. - Added BF16 Matrix-Add and Matrix-Mul support for AVX512_256 F32 kernels. - Modified the condition check in FP32 Zero point in AVX512 kernels, and fixed few bugs in Col-major Zero point evaluation and instruction usage. AMD Internal: [ CPUPL - 6748 ] * Bug Fixes in FP32 Kernels: - The current implementation lets m=1 tiny cases inside LPGEMV_TINY loop, but doesn't have the call to GEMV_M_ONE kernels. Added the m=1 path in LPGEMV_TINY loop by handling the pack A/Pack B/reorder B conditions. - Added BF16 support for BIAS, Matrix-Add and Matrix-Mul for AVX512 F32 main and GEMV kernels. - Added BF16 Downscale, BIAS, Matrix-Add and Matrix-Mul support in AVX2 GEMV_N and AVX512_256 GEMV kernels. - Added BF16 Matrix-Add and Matrix-Mul support for AVX512_256 F32 kernels. - Modified the condition check in FP32 Zero point in AVX512 kernels, and fixed few bugs in Col-major Zero point evaluation and instruction usage. AMD Internal: [ CPUPL - 6748 ] * Bug Fixes in FP32 Kernels: - The current implementation lets m=1 tiny cases inside LPGEMV_TINY loop, but doesn't have the call to GEMV_M_ONE kernels. Added the m=1 path in LPGEMV_TINY loop by handling the pack A/Pack B/reorder B conditions. - Added BF16 support for BIAS, Matrix-Add and Matrix-Mul for AVX512 F32 main and GEMV kernels. - Added BF16 Downscale, BIAS, Matrix-Add and Matrix-Mul support in AVX2 GEMV_N and AVX512_256 GEMV kernels. - Added BF16 Matrix-Add and Matrix-Mul support for AVX512_256 F32 kernels. - Modified the condition check in FP32 Zero point in AVX512 kernels, and fixed few bugs in Col-major Zero point evaluation and instruction usage. AMD Internal: [ CPUPL - 6748 ] --------- Co-authored-by: VarshaV <varshav2@amd.com>	2025-06-06 17:48:50 +05:30
Vankadari, Meghana	37efbd284e	Added 6x16 and 6xlt16 main kernels for f32 using AVX512 instructions (#38 ) * Implemented 6xlt8 AVX2 kernel for n<8 inputs * Implemented fringe kernels for 6x16 and 6xlt16 AVX512 kernels for FP32 * Implemented m-fringe kernels for 6xlt8 kernel for AVX2 * Implemented m-fringe kernels for 6xlt8 kernel for AVX2 * Added the deleted kernels and fixed bias bug AMD-Internal: SWLCSG-3556	2025-06-05 15:17:02 +05:30
V, Varsha	532eab12d3	Bug Fixes in LPGEMM for AVX512(SkyLake) machine (#24 ) * Bug Fixes in LPGEMM for AVX512(SkyLake) machine - B-matrix in bf16bf16f32obf16/f32 API is re-ordered. For machines that doesn't support BF16 instructions, the BF16 input is unre-ordered and converted to FP32 to use FP32 kernels. - For n = 1 and k = 1 sized matrices, re-ordering in BF16 is copying the matrix to the re-ordered buffer array. But the un-reordering to FP32 requires the matrix to have size multiple of 16 along n and multiple of 2 along k dimension. - The entry condition to the above has been modified for AVX512 configuration. - In bf16 API, the tiny path entry check has been modified to prevent seg fault while AOCL_ENABLE_INSTRUCTIONS=AVX2 is set in BF16 supporting machines. - Modified existing store instructions in FP32 AVX512 kernels to support execution in machines that has AVX512 support but not BF16/VNNI(SkyLake). - Added Bf16 beta and store types in FP32 avx512_256 kernels AMD Internal: [SWLCSG-3552] * Bug Fixes in LPGEMM for AVX512(SkyLake) machine - B-matrix in bf16bf16f32obf16/f32 API is re-ordered. For machines that doesn't support BF16 instructions, the BF16 input is unre-ordered and converted to FP32 to use FP32 kernels. - For n = 1 and k = 1 sized matrices, re-ordering in BF16 is copying the matrix to the re-ordered buffer array. But the un-reordering to FP32 requires the matrix to have size multiple of 16 along n and multiple of 2 along k dimension. - The entry condition to the above has been modified for AVX512 configuration. - In bf16 API, the tiny path entry check has been modified to prevent seg fault while AOCL_ENABLE_INSTRUCTIONS=AVX2 is set in BF16 supporting machines. - Modified existing store instructions in FP32 AVX512 kernels to support execution in machines that has AVX512 support but not BF16/VNNI(SkyLake). - Added Bf16 beta and store types, along with BIAS and ZP in FP32 avx512_256 kernels AMD Internal: [SWLCSG-3552] * Bug Fixes in LPGEMM for AVX512(SkyLake) machine - Support added in FP32 512_256 kerenls for : Beta, BIAS, Zero-point and BF16 store types for bf16bf16f32obf16 API execution in AVX2 mode. - B-matrix in bf16bf16f32obf16/f32 API is re-ordered. For machines that doesn't support BF16 instructions, the BF16 input is unre-ordered and converted to FP32 type to use FP32 kernels. - For n = 1 and k = 1 sized matrices, re-ordering in BF16 is copying the matrix to the re-ordered buffer array. But the un-reordering to FP32 requires the matrix to have size multiple of 16 along n and multiple of 2 along k dimension. The entry condition here has been modified for AVX512 configuration. - Fix for seg fault with AOCL_ENABLE_INSTRUCTIONS=AVX2 mode in BF16/VNNI ISA supporting configruations: - BF16 tiny path entry check has been modified to take into account arch_id to ensure improper entry into the tiny kernel. - The store in BF16->FP32 col-major for m = 1 conditions were updated to correct storage pattern, - BF16 beta load macro was modified to account for data in unaligned memory. - Modified existing store instructions in FP32 AVX512 kernels to support execution in machines that has AVX512 support but not BF16/VNNI(SkyLake) AMD Internal: [SWLCSG-3552] --------- Co-authored-by: VarshaV <varshav2@amd.com>	2025-05-30 17:22:49 +05:30
Negi, Deepak	121d81df16	Implemented GEMV kernel for m=1 case. (#5 ) * Implemented GEMV kernel for m=1 case. Description: - Added a new GEMV kernel for AVX2 where m=1. - Added a new GEMV kernel for AVX512 with ymm registers where m=1.	2025-05-13 16:33:04 +05:30
Meghana Vankadari	8557e2f7b9	Implemented GEMV for n=1 case using 32 YMM registers Details: - This implementation is picked form cntx when GEMM is invoked on machines that support AVX512 instructions by forcing the AVX2 path using AOCL_ENABLE_INSTRUCTIONS=AVX2 during run-time. - This implementation uses MR=16 for GEMV. AMD-Internal: [SWLCSG-3519] Change-Id: I8598ce6b05c3d5a96c764d96089171570fbb9e1a	2025-05-05 05:31:13 -04:00
Meghana Vankadari	21aa63eca1	Implemented AVX2 based GEMV for n=1 case. - Added a new GEMV kernel with MR = 8 which will be used for cases where n=1. - Modified GEMM and GEMV framework to choose right GEMV kernel based on compile-time and run-time architecture parameters. This had to be done since GEMV kernels are not stored-in/retrieved-from the cntx. - Added a pack kernel that packs A matrix from col-major to row-major using AVX2 instructions. AMD-Internal: [SWLCSG-3519] Change-Id: Ibf7a8121d0bde37660eac58a160c5b9c9ebd2b5c	2025-05-05 08:56:22 +00:00
Meghana Vankadari	4745cf876e	Implemented a new set of kernels for f32 using 32 YMM regs Details: - These kernels are picked from cntx when GEMM is invoked on machines that support AVX512 instructions by forcing the AVX2 path using AOCL_ENABLE_INSTRUCTIONS=AVX2 during run-time. - This path uses the same blocksizes and pack kernels as AVX512 path. - GEMV is disabled currently as AVX2 kernels for GEMV are not implemented. AMD-Internal: [SWLCSG-3519] Change-Id: I75401fac48478fe99edb8e71fa44d36dd7513ae5	2025-04-23 12:02:01 +00:00
Deepak Negi	48c7452b08	Beta and Downscale support for F32 AVX-512 kernels Description - To enable AVX512 VNNI support without native BF16 in BF16 kernels, the BF16 C_type is converted to F32 for computation and then cast back to BF16 before storing the result. - Added support for handling BF16 zero-point values of BF16 type. - Added a condition to disable the tiny path for the BF16 code path where native BF16 is not supported. AMD Internal : [CPUPL-6627] Change-Id: I1e0cfefd24c5ffbcc95db73e7f5784a957c79ab9	2025-04-23 06:12:14 -05:00
Arnav Sharma	87c9230cac	Bugfix: Disable A Packing for FP32 RD kernels and Post-Ops Fix - For single-threaded configuration of BLIS, packing of A and B matrices are enabled by default. But, packing of A is only supported for RV kernels where elements from matrix A are being broadcasted. Since elements are being loaded in RD kernels, packing of A results in failures. Hence, disabled packing of matrix A for RD kernels. - Fixed the issue where c_i index pointer was incorrectly being reset when exceeding MC block thus, resulting in failures for certain Post-Ops. - Fixed the FP32 reoder case were for n == 1 and rs_b == 1 condition, it was incorrectly using sizeof(BLIS_FLOAT) instead of sizeof(float). AMD-Internal: [SWLCSG-3497] Change-Id: I6d18afa996c253d79f666ea9789270bb59b629dd	2025-04-18 14:31:03 +05:30
Arnav Sharma	267aae80ea	Added Post-Ops Support for F32 RD Kernels - Support for Post-Ops has been added for all F32 RD AVX512 and AVX2 kernels. AMD-Internal: [SWLCSG-3497] Change-Id: Ia2967417303d8278c547957878d93c42c887109e	2025-04-11 05:25:30 -04:00
Arnav Sharma	c68c258fad	Added AVX512 and AVX2 FP32 RD Kernels - Added FP32 RD (dot-product) kernels for both, AVX512 and AVX2 ISAs. - The FP32 AVX512 primary RD kernel has blocking of dimensions 6x64 (MRxNR) whereas it is 6x16 (MRxNR) for the AVX2 primary RD kernel. - Updatd f32 framework to accomodate rd kernels in case of B trans with thresholds - Updated data gen python script TODO: - Post-Ops not yet supported. Change-Id: Ibf282741f58a1446321273d5b8044db993f23714	2025-04-05 20:16:51 -05:00
Hari Govind S	8998839c71	Optimisation of DGEMV Transpose Case for unit stride - Included a new code section to handle input having non-unit strided y vector for dgemv transpose case. Removed the same from the respective kernels to avoid repeated branching caused by condition checks within the 'for' loop. - The condition check for beta is equal to zero in the primary kernels are moved outside the for loop to avoid repeated branching. - The '_mm512_reduce_pd' operations in the primary kernel is replaced by a series of operations to reduce the number of instructions required to reduce the 8 registers. - Changing naming convention for DGEMV transpose kernels. - Modified unit kernel test to avoid y increment for dgemv tranpose kernels during the test. AMD-Internal: [CPUPL-6565] Change-Id: I1ac516d6b8f156ac53ac9f6eb18badd50e152e05	2025-03-06 05:15:58 -05:00
Arnav Sharma	b4c1026ec2	Added Support for General Stride in DGEMV - Updated the bli_dgemv_zen_ref( ... ) kernel to support general stride. - Since the latest dgemv kernels don't support general stride, added checks to invoke bli_dgemv_zen_ref( ... ) when A matrix has a general stride. - Thanks to Vignesh Balasubramanian <vignesh.balasubramanian@amd.com> for finding this issue. AMD-Internal: [CPUPL-6492] Change-Id: Ia987ce7674cb26cb32eea4a6e9bd6623f2027328	2025-02-27 12:47:21 -05:00
Mithun Mohan	7394aafd1e	New A packing kernels for F32 API in LPGEMM. -New packing kernels for A matrix, both based on AVX512 and AVX2 ISA, for both row and column major storage are added as part of this change. Dependency on haswell A packing kernels are removed by this. -Tiny GEMM thresholds are further tuned for BF16 and F32 APIs. AMD-Internal: [SWLCSG-3380, SWLCSG-3415] Change-Id: I7330defacbacc9d07037ce1baf4a441f941e59be	2025-02-26 05:23:35 +00:00
varshav	8a69141294	Bug fix in BF16-F32 supported AVX2 Kernels - Bug fix in Matrix Mul post op. - Updated the config in AVX512_VNNI_BF16 context to work in AVX2 kernels Change-Id: I25980508facc38606596402dba4cfce88f4eb173	2025-02-25 14:42:45 +00:00
varshav	a0005c60ce	Add col-major pack kernels and BF16 output support in F32 AVX-2 kernels. - Added column major pack kernels, which will transpose and store the BF16 matrix input to F32 input matrix - Added BF16 Zero point Downscale support to F32 main and fringe kernels. - Updated Matrix Add and Matrix Mul post-ops in f32-AVX2 main and fringe kernels to support BF16 input. - Modified the f32 tiny kernels loop to update the buf_downscale parameter. - Modified bf16bf16f32obf16 framework to work with AVX-2 system. - Added wrapper in bf16 5-Loop to call the corresponding AVX-2/AVX-512 5 Loop functions. - Bug fixes in the f32-AVX2 kernels BIAS post-ops. - Bug fixes in the Convert function, and the bf16 5-loop for multi-threaded inputs. AMD-Internal:[SWLCSG-3281 , CPUPL-6447] Change-Id: I4191fbe6f79119410c2328cd61d9b4d87b7a2bcd	2025-02-24 09:51:12 +05:30
varshav2	ef04388a44	Added AVX2 support for BF16 kernels: Row major - Currently the BF16 kernels uses the AVX512 VNNI instructions. In order to support AVX2 kernels, the BF16 input has to be converted to F32 and then the F32 kernels has to be executed. - Added un-pack function for the B-Matrix, which does the unpacking of the Re-ordered BF16 B-Matrix and converts it to Float. - Added a kernel, to convert the matrix data from Bf16 to F32 for the give input. - Added a new path to the BF16 5LOOP to work with the BF16 data, where the packed/unpacked A matrix is converted from BF16 to F32. The packed B matrix is converted from BF16 to F32 and the re-ordered B matrix is unre-ordered and converted to F32 before feeding to the F32 micro kernels. - Removed AVX512 condition checks in BF16 code path. - Added the Re-order reference code path to support BF16 AVX2. - Currently the F32 AVX-2 kernels supports only F32 BIAS support. Added BF16 support for BIAS post-op in F32 AVX2 kernels. - Bug fix in the test input generation script. AMD Internal : [SWLCSG - 3281] Change-Id: I1f9d59bfae4d874bf9fdab9bcfec5da91eadb0fb	2025-02-10 08:18:52 -05:00
Mithun Mohan	bffa92ec93	Deprecate S16 LPGEMM APIs. -The following S16 APIs are removed: 1. aocl_gemm_u8s8s16os16 2. aocl_gemm_u8s8s16os8 3. aocl_gemm_u8s8s16ou8 4. aocl_gemm_s8s8s16os16 5. aocl_gemm_s8s8s16os8 along with the associated reorder APIs and corresponding framework elements. AMD-Internal: [CPUPL-6412] Change-Id: I251f8b02a4cba5110615ddeb977d86f5c949363b	2025-02-07 11:43:28 +00:00
Edward Smyth	1f0fb05277	Code cleanup: Copyright notices (2) More changes to standardize copyright formatting and correct years for some files modified in recent commits. AMD-Internal: [CPUPL-5895] Change-Id: Ie95d599710c1e0605f14bbf71467ca5f5352af12	2025-02-07 05:41:44 -05:00
Hari Govind S	3d2653f1ab	DDOTV Optimization for ZEN3 Architecture - Reduced the blocking size of 'bli_ddotv_zen_int10' kernel from 40 elements to 20 elements for better utilization of vector registers - Replaced redundant 'for' loops in 'bli_ddotv_zen_int10' kernel with 'if' conditions to handle reminder iterations. As only a single iteration is used when reminder is less than the primary unroll factor. - Added a conditional check to invoke the vectorized DDOTV kernels directly(fast-path), without incurring any additional framework overhead. - The fast-path is taken when the input size is ideal for single-threaded execution. Thus, we avoid the call to bli_nthreads_l1() function to set the ideal number of threads. - Updated getestsuite ukr tests for 'bli_ddotv_zen_int10' kernel. AMD-Internal: [CPUPL-4877] Change-Id: If43f0fcff1c5b1563ad233005717398b5b6fb8f2	2025-02-04 06:01:04 -05:00
Deepak Negi	4ade159800	Buffer scale support for matrix add and matrix mul post-ops in f32 API. Description: 1. Support has been added to scale buffer values using both scalar and vector scale factors before matrix add or matrix mul post-ops. AMD-Internal: CPUPL-6340 Change-Id: Ie023d5963689897509ef3d5784c3592791e57125	2025-01-30 07:09:18 -05:00
Vignesh Balasubramanian	fb6dcc4edb	Support for Tiny-GEMM interface(ZGEMM) - As part of AOCL-BLAS, there exists a set of vectorized SUP kernels for GEMM, that are performant when invoked in a bare-metal fashion. - Designed a macro-based interface for handling tiny sizes in GEMM, that would utilize there kernels. This is currently instantiated for 'Z' datatype(double-precision complex). - Design breakdown : - Tiny path requires the usage of AVX2 and/or AVX512 SUP kernels, based on the micro-architecture. The decision logic for invoking tiny-path is specific to the micro-architecture. These thresholds are defined in their respective configuration directories(header files). - List of AVX2/AVX512 SUP kernels(lookup table), and their lookup functions are defined in the base-architecture from which the support starts. Since we need to support backward compatibility when defining the lookup table/functions, they are present in the kernels folder(base-architecture). - Defined a new type to be used to create the lookup table and its entries. This type holds the kernel pointer, blocking dimensions and the storage preference. - This design would only require the appropriate thresholds and the associated lookup table to be defined for the other datatypes and micro-architecture support. Thus, is it extensible. - NOTE : The SUP kernels that are listed for Tiny GEMM are m-var kernels. Thus, the blocking in framework is done accordingly. In case of adding the support for n-var, the variant information could be encoded in the object definition. - Added test-cases to validate the interface for functionality(API level tests). Also added exception value tests, which have been disabled due to the SUP kernel optimizations. AMD-Internal: [CPUPL-6040][CPUPL-6018][CPUPL-5319][CPUPL-3799] Change-Id: I84f734f8e683c90efa63f2fa79d2c03484e07956	2025-01-24 12:59:26 -05:00
Hari Govind S	c149b5a98b	Optimisation of DCOPY kernel in zen3 - Using 'if' condition instead of 'for'loop to handle fringe cases. 'for' loop is redundant for handling reminder iterations as only a single iteration is used when reminder is less than primary unroll factor. AMD-Internal: [CPUPL-5594] Change-Id: I8cebc037742ee47961869e22e2471e550fcd99e9	2025-01-24 05:55:59 -05:00
Hari Govind S	349fc47ec5	DGEMV Optimizations for TRANSPOSE Cases - Developed new AVX512 DGEMV kernels for Zen4/5 architectures and AVX2 kernels for Zen1/2/3 architectures. These kernels are written from the ground up and are independent of fused kernels. - The DGEMV primary kernel processes the calculation in chunks of 8 columns. Fringe columns (sizes 1 to 7) are handled by fringe kernels, which are invoked by the primary kernel as needed. - Implemented the kernels by computing the dot product of matrix A columns with vector x in chunks of 32 elements, storing the results in accumulator registers. Fringe elements are handled in chunks of 16, 8, etc. The data in the accumulator registers is then reduced and added to vector y. AMD-Internal: [CPUPL-5835] Change-Id: I5cb9eb1330db095931586a7028fd7676fbbecc61	2025-01-24 00:38:34 -05:00
Vignesh Balasubramanian	a80436ab21	Standardizing the EVT compliance of {S/D}AMAXV API - Updated the existing AVX2 {S/D}AMAXV kernels to comply to the standard when having exception values. This makes it exhibit the same behaviour as it AVX512 variants. Provided additional optimizations with loop unrolling. - Removed redundant early return checks inside the kernels, since they have been abstracted to a higher layer. - Updated the unit-tests(micro-kernel) and exception value tests for appropriate code-coverage. Also re-enabled the exception value tests. AMD-Internal: [CPUPL-4745] Change-Id: I36c793220bd4977a00281af9737c51cd1e5c60d9	2025-01-13 06:56:31 -05:00
Edward Smyth	97ede96ed4	Correct duplicate object file names Some kernel file names were the same for different sub-configurations, which could result in duplicate copies of the same object being archived depending upon the order of (re-)compiling the source files. Rename the files to be specific to each sub-configuration to avoid this problem. AMD-Internal: [CPUPL-5895] Change-Id: I182ac706e04a364f1df20fd0fb5b633eb10eeafb	2025-01-10 06:03:36 -05:00
Meghana Vankadari	852cdc6a9a	Implemented batch_matmul for f32 & int8 datatypes Details: - The batch matmul performs a series of matmuls, processing more than one GEMM problem at once. - Introduced a new parameter called batch_size for the user to indicate number of GEMM problems in a batch/group. - This operation supports processing GEMM problems with different parameters including dims,post-ops,stor-schemes etc., - This operation is optimized for problems where all the GEMMs in a batch are of same size and shape. - For now, the threads are distributed among different GEMM problems equally irrespective of their dimensions which leads to better performance for batches with identical GEMMs but performs sub-optimally for batches with non-identical GEMMs. - Optimizations for batches with non-identical GEMMs is in progress. - Added bench and input files for batch_matmul. - Added logger functionality for batch_matmul APIs. AMD-Internal: [SWLCSG-2944] Change-Id: I83e26c1f30a5dd5a31139f6706ac74be0aa6bd9a	2025-01-10 04:10:53 -05:00
Meghana Vankadari	051c9ac7a2	Bug fixes in F32 and INT8 APIs Details: - Fixed few bugs in downscale post-op for f32 datatype. - Fixed a bug in setting strides of packB buffer in int8 APIs. Change-Id: Idb3019cc4593eace3bd5475dd1463dea32dbe75c	2025-01-09 04:07:26 -05:00
harsdave	54b46ec1ed	Enhance 24x8 DGEMM SUP/Tiny Kernel Performance with Optimized Loops and Edge Kernels This patch introduces comprehensive optimizations to the DGEMM kernel, focusing on loop efficiency and edge kernel performance. The following technical improvements have been implemented: 1. IR Loop Optimization: - The IR loop has been re-implemented in hand-written assembly to eliminate the overhead associated with `begin_asm` and `end_asm` calls, resulting in more efficient execution. 2. JR Loop Integration: - The JR loop is now incorporated into the micro kernel. This integration avoids the repetitive overhead of stack frame management for each JR iteration, thereby enhancing loop performance. 3. Kernel Decomposition Strategy: - The m dimension is decomposed into specific sizes: 20, 18, 17, 16, 12, 11, 10, 9, 8, 4, 2, and 1. - For remaining cases, masked variants of edge kernels are utilized to handle the decomposition efficiently. 1. Interleaved Scaling by Alpha: - Scaling by the alpha factor is interleaved with load instructions to optimize the instruction pipeline and reduce latency. 2. Efficient Mask Preparation: - Masks are prepared within inline assembly code only at points where masked load-store operations are necessary, minimizing unnecessary overhead. 3. Broadcast Instruction Optimization: - In edge kernels where each FMA (Fused Multiply-Add) operation requires a broadcast without subsequent reuse, the broadcast instruction is replaced with `mem_1to8`. - This allows the compiler to optimize by assigning separate vector registers for broadcasting, thus avoiding dependency chains and improving execution efficiency. 4. C Matrix Update Optimization: - During the update of the C matrix in edge kernels, columns are pre-loaded into multiple vector registers. This approach breaks dependency chains during FMA operations following the scaling by alpha, thereby mitigating performance bottlenecks and enhancing throughput. These optimizations collectively improve the performance of the DGEMM kernel, particularly in handling edge cases and reducing overhead in critical loops. The changes are expected to yield significant performance gains in matrix multiplication operations. This patch also involves changes for tiny gemm interface. A light interface for calling kernels and removing calls to avx2 dgemm kernels as we use avx512 dgemm kernels for all the sizes for zen4 and zen5. For zen4 and zen5 when A matrix transposed(CRC, RRC), tiny kernel does not have the support to handle such inputs and thus such inputs are routed to gemm_small path. AMD-Internal: [CPUPL-6054] Change-Id: I57b430f9969ca39aa111b54fa169e4225b900c4a	2024-12-13 00:03:00 -05:00
Arnav Sharma	25e59fcbb9	DGEMV Optimizations for NO_TRANSPOSE Cases - AVX512 specific DGEMV native kernels are added for Zen4/5 architectures to handle the NO_TRANSPOSE cases and are independent of the AXPYF fused kernels. - The following set of kernels biased towards the n-dimension perform beta scaling of y vector within the kernel itself and handle cases where n is less than 5: - bli_dgemv_n_zen_int_32x8n_avx512( ... ) - bli_dgemv_n_zen_int_32x4n_avx512( ... ) - bli_dgemv_n_zen_int_32x2n_avx512( ... ) - bli_dgemv_n_zen_int_32x1n_avx512( ... ) - The bli_dgemv_n_zen_int_16mx8_avx512( ... ) is biased towards the m-dimension and for this kernel beta scaling is handled beforehand within the framework. - Added unit-tests for the new kernels. - AVX2 path for Zen/2/3 architectures still follows the old approach of using fused kernel, namely AXPYF, to perform the GEMV operation. AMD-Internal: [CPUPL-5560] Change-Id: I22bc2a865cd28b9cdcb383e17d1ff38bdd28de79	2024-12-12 10:26:50 -05:00
Deepak Negi	60a8c71a1a	Sigmoid and Tanh post-operation support for int8 API's. Description: Implemented sigmoid, tanh as fused post-ops in aocl_gemm_<s8\|u8>s8<s32\|s16>o<s8\|u8\|s32> API's Sigmoid(x) = 1/1+e^(-x) Tanh(x) = (1-e^(-2x))/(1+e^(2x)) Updated bench_lpgemm to recognize sigmod, tanh as options for post-ops from bench_input and verified. AMD-Internal: [SWLCSG-3178] Change-Id: I9df3aab02222f728ff9d1f292c7bc549f30176f0	2024-11-15 05:36:31 -05:00
Deepak Negi	146f3b2eb2	Sigmoid and Tanh post-operation support for f32 API. Description: Implemented sigmoid, tanh as fused post-ops in aocl_gemm_f32f32f32of32 API's Sigmoid(x) = 1/1+e^(-x) Tanh(x) = (1-e^(-2x))/(1+e^(2x)) Updated bench_lpgemm to recognize sigmod, tanh as options for post-ops from bench_input and verified. AMD-Internal: [SWLCSG-3178] Change-Id: Iac0a907f6dea1d9cb82d9fd8716bfdbf1c33921d	2024-11-15 04:20:20 -04:00
Mithun Mohan	097cda9f9e	Adding support for AOCL_ENABLE_INSTRUCTIONS for f32 LPGEMM API. -Currently lpgemm sets the context (block sizes and micro-kernels) based on the ISA of the machine it is being executed on. However this approach does not give the flexibility to select a different context at runtime. In order to enable runtime selection of context, the context initialization is modified to read the AOCL_ENABLE_INSTRUCTIONS env variable and set the context based on the same. As part of this commit, only f32 context selection is enabled. -Bug fixes in scale ops in f32 micro-kernels and GEMV path selection. -Added vectorized f32 packing kernels for NR=16(AVX2) and NR=64(AVX512). This is only for B matrix and helps remove dependency of f32 lpgemm api on the BLIS packing framework. AMD Internal: [CPUPL-5959] Change-Id: I4b459aaf33c54423952f89905ba43cf119ce20f6	2024-10-30 08:52:22 +00:00
varshav2	dabfdf484a	Add Scale post-op for F32 API - Implemented the Scale post-op for the F32 API for all kernels - f32_scale = (f32 * scale_factor) + offset - Added the bench inputs Change-Id: Ib0f25f870eafe695d8b2a2c434c8cb3ec4f7db4c	2024-10-21 06:08:31 -04:00
Edward Smyth	a07e041b1f	SCALV alpha=zero BLAS compliance SCALV is used directly by BLAS, CBLAS and BLIS scal{v} APIs but also within many other APIs to handle special cases. In general it is preferred to use SETV when alpha=0, but BLAS and CBLAS continue to multiple all vector element by alpha. This has different behaviour for propagating NaNs or Infs. Changes in this commit: - Standardize early returns from SCALV reference and optimized kernels. - User supplied N<0 is handled at the top level API layer. Use negative values of N in kernel calls to signify that SETV should _not_ be used when alpha=0. This should only be required in SCALV. - Include serial threshold in zdscal (as in dscal) to reduce overhead for small problem sizes. - Code tidying to make different variants more consistent. - More standardization of tests in SCALV gtestsuite programs. - Remove scalv_extreme_cases.cpp as it is now redundant. AMD-Internal: [CPUPL-4415] Change-Id: I42e98875ceaea224cc98d0cdfe0133c9abc3edae	2024-09-16 07:10:28 -04:00
Vignesh Balasubramanian	ed682a429d	Fixed compiler warnings due to prefetch in AVX2 and AVX512 kernels - Added explicit typecast to the pointers that are passed to the _mm_prefetch( ... ) intrinsic, to avoid compiler warnings. AMD-Internal: [CPUPL-4415] Change-Id: I1c1398b7b5abe81848d33cb6df107f7f077588ea	2024-09-12 11:37:06 +05:30
Chandrashekara K R	2ff0125f11	CMake: Enabled ADDON(aocl_gemm) feature for Windows. 1. Updated datatype from __int64_t to int64_t. Since __int64_t was not defined for Windows 2. Updated CMake build system to build lpgemm on windows Change-Id: I5fc5ed93ecc54e4a9931b7b40b790d37c7ead4b8	2024-08-23 07:05:28 -04:00
Vignesh Balasubramanian	fd47d0aebb	Fixing compiler warnings when exporting internal BLIS kernels - Added the attribute to export symbols, in the header file that contains the L1 kernel declarations. This attribute was previously added as part of the kernel definitions. AMD-Internal: [CPUPL-4415] Change-Id: I375246f47d53c220f885644f9b75c7d7991ae710	2024-08-12 12:30:51 +05:30
Meghana Vankadari	5514c7a75f	Added LPGEMV(n=1) kernels for s8s8s32os32\|s8 and s8s8s16os16\|s8 APIs - When n=1, reorder of B matrix is avoided to efficiently process data. A dot-product based kernel is implemented to perform gemv when n==1. AMD-Internal: [SWLCSG-2354] Change-Id: I6b73dfddd9a15e7b914d031646a1d913a7ab4761	2024-08-09 06:17:52 -04:00
Edward Smyth	7fff7b4026	Code cleanup: Miscellaneous fixes - Delete unused cmake files. - Add guards around call to bli_cpuid_is_avx2fma3_supported in frame/3/bli_l3_sup.c, currently assumes that non-x86 platforms will not use bli_gemmtsup. - Correct variable in frame/base/bli_arch.c on non-x86 builds. - Add guards around omp pragma to avoid possible gcc compiler warning in kernels/zen/2/bli_gemv_zen_int_4.c. - Add missing registers in clobber list in kernels/zen4/1/bli_dotv_zen_int_avx512.c. - Add gtestsuite ERS_IIT tests for TRMV, copied from TRSV. - Correct calls to cblas_{c,z}swap in gtestsuite. - Correct test name in ddotxf gtestsuite program. AMD-Internal: [CPUPL-4415] Change-Id: I69ad56390017676cc609b4d3aba3244a2df6a6b5	2024-08-06 06:56:01 -04:00
Edward Smyth	89f52a6df5	Code cleanup: spelling corrections Corrections for spelling and other mistakes in code comments and doc files. AMD-Internal: [CPUPL-4500] Change-Id: I33e28932b0e26bbed850c55602dee12fd002da7f	2024-08-05 16:18:51 -04:00
Edward Smyth	82bdf7c8c7	Code cleanup: Copyright notices - Standardize formatting (spacing etc). - Add full copyright to cmake files (excluding .json) - Correct copyright and disclaimer text for frame and zen, skx and a couple of other kernels to cover all contributors, as is commonly used in other files. - Fixed some typos and missing lines in copyright statements. AMD-Internal: [CPUPL-4415] Change-Id: Ib248bb6033c4d0b408773cf0e2a2cda6c2a74371	2024-08-05 15:35:08 -04:00
Edward Smyth	591a3a7395	Code cleanup: file formats and permissions - Remove execute file permission from source and make files. - dos2unix conversion. - Add missing eol at end of files. Also update .gitignore to not exclude build directory but to exclude any build_* created by cmake builds. AMD-Internal: [CPUPL-4415] Change-Id: I5403290d49fe212659a8015d5e94281fe41eb124	2024-08-05 11:52:33 -04:00
mkadavil	9f5fec7713	Matrix MUL op support in element wise operations API for bfloat16. -Matrix MUL op support added in main as well as fringe bfloat16 element wise operations kernels. -Benchmarking/testing framework for the same is added. -Fixed issues in setting up post-ops node index. AMD Internal: [SWLCSG-2947, SWLCSG-2953] Change-Id: Iba7561a6a60df41211efbf06fab1b4900207bcf8	2024-08-05 08:29:42 +05:30
Deepak Negi	80bf6249f0	Matrix MUL post-operation support for float(bf16\|f32) LPGEMM APIs. This post-operation computes C = (betaC + alphaAB) D, where D is a matrix with dimensions and data type the same as that of C matrix. AMD-Internal: [SWLCSG-2953] Change-Id: Id4df2ca76a8f696cb16edbd02c25f621f9a828fd	2024-08-05 08:25:32 -04:00
Arnav Sharma	0a5c057475	DGEMV Optimizations for Tiny Sizes - Added reference kernel for dgemv that handles computation for tiny sizes (m < 8 && n < 8). - The reference kernel, bli_dgemv_zen_ref( ... ), supports both row/column storage schemes as well as transpose and no transpose cases. - Added additional unit-tests for functional verification. AMD-Internal: [CPUPL-5098] Change-Id: I66fdf0a40e90bdb3fed40152c45ab28a17a87ada	2024-08-05 12:19:42 +05:30
Moripalli Chitra	8b486e8d14	Added new decision logic to choose between 6x8 dgemm kernel vs 24x8 kernel. The decision is based on the values of "m, n and k". Change-Id: I307ff002797ccef5bd61106b808cecb069b91fd6	2024-08-02 14:18:58 +05:30
Vignesh Balasubramanian	f23b8e636b	AVX2 and AVX512 optimizations for DAXPYV - Removed some of the unrolling factors that affected the performance of AVX2 DAXPYV kernel. In addition to improving the current performance on sizes compatible to single-threaded runs, this will now perform better for tiny sizes as well since the overhead to reach the computation is less. - Updated the vector partitioning logic, by using bli_thread_range_sub( ... ), which ensures that there is no false sharing among multiple threads. - Updated the AOCL-DYNAMIC logic for the API, to include thresholds or zen4 and zen5 micro-architectures. AMD-Internal: [CPUPL-5514] Change-Id: Iee9edddac685334213cd6694421ab3df3547e930	2024-07-31 09:24:36 -04:00

1 2 3 4 5 ...

366 Commits