amd/blis - blis - Public git mirror

amd/blis

mirror of https://github.com/amd/blis.git synced 2026-05-24 18:34:40 +00:00

Author	SHA1	Message	Date
Meghana Vankadari	d5b4d3aa5e	Fixing control flow in aocl_gemm_bf16s4f32of32\|bf16 - Fixed framework of bf16s4f32of32 API to correct pointer updations. - Modified pre_op structure to exclude pre-op-offset. Now offset is passed as a separate parameter to the scale-pack functions. - Fixed work-distribution among threads in MT scenario. - Added Blocksizes and kernel-pointers and verified functionality for the new API. AMD-Internal: [SWLCSG-2943] Change-Id: I58fece240d62c798c880a2b2b7fa64e560cc753d	2024-07-29 05:12:09 -04:00
mkadavil	ec8c39541e	Test/benchmark framework updates to test WOQ workflow. -To enable Weight-only-Quantization (WOQ) workflow, new LPGEMM APIs have been developed where data types are A:bf16, B:int4 and C:f32/bf16. The testing and benchmarking framework for the same are added. AMD-Internal: [SWLCSG-2943] Change-Id: Icdc1d60819a23dd9f41382499d1a3c055c5edc17	2024-07-25 06:44:37 +05:30
mkadavil	d37c91dffa	Quantization (scale + zero point) support for BF16 LPGEMM api. -Quantization of f32 to bf16 (bf16 = (f32 * scale_factor) + zero_point) instead of just type conversion in aocl_gemm_bf16bf16f32obf16. -Support for multiple scale/sum/matrix_add/bias post-ops in a single LPGEMM api call. -Post-ops mask related fixes in lpgemv kernels . -Additional scale post-ops sanity checks. AMD-Internal: [SWLCSG-2945] Change-Id: I3b35cc413c176bb50bfdbd6acd4839a5ba7e94bb	2024-07-18 05:32:51 -04:00
srikanth pogula	1d7f6d414f	Bench APPs - change in Print statement for more params >Made changes in the print statements in bench files to print all the params of the individual APIs > Ex : removing tab & adding Func param "Dt\t n\t incx\t incy\t gflops\n" --> "Func Dt n incx incy gflops\n" > Ex : adding func, incx, incy params "dt_ch, n, alpha_r, alpha_i, beta_r, beta_i, gflops" --> "tmp, dt_ch, n, alpha_r, alpha_i, incx, beta_r, beta_i, incy, gflops" Change-Id: Ib5d151d7472d3f88c13a85a615a447dfa5e6b528	2024-07-11 02:04:19 -04:00
Nallani Bhaskar	d5133e4363	Fixed linking issue with bli_print_msg function Description: In recent changes bli_print_msg is used in lpgemm test application file bench_lpgemm.c for printing error message. bli_print_msg is a blis library function which is not exported for the usage of applications, because of which linking failed when blis shared library is used to build. Updated bli_print_msg with printf in the bench_lpgemm.c AMD Internal: CPUPL-5326 Change-Id: I021849baa6881bd997013e42013db1c5c711627f	2024-07-09 23:18:53 +05:30
Varaganti, Kiran	2ac24d1f9c	Avoided Extra copy of "c" matrix Initailized c_save instead of 'c" and then removed copying c to c_save. Because at the start every n_repeats iteration we are copying back c_save to c. Therefore if we initialize c_save, we can avoid extra copy of "c" to c_save before calling GEMM. For very large sizes matrix initialization takes considerable amount of time. This can be reduced now. Change-Id: I2c6ffe169e991607314897cb0c1fbfc0d74ef179	2024-07-09 00:54:03 -04:00
Meghana Vankadari	4e6fa17c08	Bug fix in LPGEMV for INT8 APIs Details: - Corrected the usage of vpdpbusd instruction in GEMV implementation for INT8 APIs. - Modified bench to fill matrices with values ranging between -5 and +5 whenever the datatype is a signed integer. Change-Id: I457462b888b667d8a34c53de762e9b4aee784ecc	2024-06-27 04:22:04 +05:30
mkadavil	a5c4a8c7e0	Int4 B matrix reordering support in LPGEMM. Support for reordering B matrix of datatype int4 as per the pack schema requirements of u8s8s32 kernel. Vectorized int4_t -> int8_t conversion implemented via leveraging the vpmultishiftqb instruction. The reordered B matrix will then be used in the u8s8s32o<s32\|s8> api. AMD-Internal: [SWLCSG-2390] Change-Id: I3a8f8aba30cac0c4828a31f1d27fa1b45ea07bba	2024-06-24 07:55:34 -04:00
Chandrashekara K R	fa75ce725e	CMake: Added logic to link openmp library given through OpenMP_libomp_LIBRARY cmake variable on linux. Enabled command line option to link libiomp5.so or libomp.so or libgomp.so libraries using cmake. Eg:- -DOpenMP_libomp_LIBRARY=<path to openmp library including library name>. If we not set above variable, by default openmp library will be libomp.so for clang and libgomp.so for gcc compiler. Change-Id: I5bffa10ff8351f5d10f0d543cbdf55aa16c84c90	2024-06-10 04:41:23 -04:00
mkadavil	cd032225ca	BF16 bias support for bf16bf16f32ob16. -As it stands the bf16bf16f32ob16 API expects bias array to be of type float. However actual use case requires the usage of bias array of bf16 type. The bf16 micro-kernels are updated to work with bf16 bias array by upscaling it to float type and then using it in the post-ops workflow. -Corrected register usage in bf16 JIT generator for bf16bf16f32ob16 API when k > KC. AMD-Internal: [SWLCSG-2604] Change-Id: I404e566ff59d1f3730b569eb8bef865cb7a3b4a1	2024-05-23 04:48:20 +05:30
Nallani Bhaskar	29db6eb42b	Added transB in all AVX512 based int8 API's Description: --Added support for tranB in u8s8s32o<s32\|s8> and s8s8s32o<s32\|s8> API's --Updated the bench_lpgemm by adding options to support transpose of B matrix --Updated data_gen_script.py in lpgemm bench according to latest input format. AMD-Internal: [SWLCSG-2582] Change-Id: I4a05cc390ae11440d6ff86da281dbafbeb907048	2024-05-23 03:46:13 +05:30
mkadavil	118e955a22	SWISH post-op support for all LPGEMM APIs. SWISH post-op computes swish(x) = x / (1 + exp(-1 * alpha * x)). SiLU = SWISH with alpha = 1. AMD-Internal: [SWLCSG-2387] Change-Id: I55f50c74a8583a515f7ea58fa0878ccbcdd6cc26	2024-05-06 06:05:11 -04:00
Vignesh Balasubramanian	1b7980a38d	Added support to benchmark AXPYV APIs - Implemented the feature to benchmark ?AXPYV APIs for the supported datatypes. The feature allows to benchmark BLAS, CBLAS or the native BLIS API, based on the macro definition. - Added a sample input file to provide examples to benchmark AXPYV for all its datatype supports. - Updated the sample input file for SCALV to provide examples to benchmark all of its datatype supports. AMD-Internal: [CPUPL-4805] Change-Id: I550920e3a57fcc2e4900e9e698330d8b8595bdee	2024-04-08 00:06:54 -04:00
Arnav Sharma	f71495a135	Support for DOTC in DOTV Bench and DTL updates - Added support for ?DOTC in bench. - Updated DTL to accept conjx as a parameter: - 'N', i.e., no conjugate for DOTU - 'C', i.e., conjugate for DOTC - Updated DTL calls in the interface with respective values of conjx. AMD-Internal: [CPUPL-4804] Change-Id: I447b19a6273566c6021c1721ce173bac4a59142c	2024-04-04 12:27:53 +05:30
jagar	bd80488af1	CMake: Update code to support blastest for ILP64 on windows Change-Id: I8e87ee073ffcb893fbcc7c9580add217ae347449	2024-03-27 12:02:26 -04:00
Eleni Vlachopoulou	020b9ff7f0	CMake: Enable builds for both static and shared builds for Linux. - Added BUILD_STATIC_LIBS option which is on by default, only on Linux. - Added TEST_WITH_SHARED option which is off by default, only on Linux. - If only shared or static lib is being built, that's the one that will be used for testing. - If both are being built, TEST_WITH_SHARED determins which library wil be used for testing. - Set linux workflows so that they build both static and shared libs, and use linux-static and linux-shared to denote which one should be used for testing. - Set -fPIC for both static and shared builds to fix issues faced when building blis using AOCC 4.0.0 and gtestsuite using gcc 9.4.0. AMD-Internal: [CPUPL-2748] Change-Id: I4227bab97ff31ecddfe218e18499f33b4e4ee63e	2024-03-14 10:32:51 -04:00
jagar	e2de45b454	CMake:Added support for ADDON(aocl_gemm) on Windows CMakelists.txt is updated to support aocl_gemm on windows. On windows, BLIS library(blis+aocl_gemm) is built successfully only with AOCC Compiler. (Clang has an issue with optimizing VNNI instructions). $cmake .. -DENABLE_ADDON="aocl_gemm" .... AMD-Internal: [CPUPL-2748] Change-Id: I9620878ab6934233fadc9ddc5d5e82ad85be9209	2024-03-14 07:57:02 -04:00
Vignesh Balasubramanian	d1a6517642	Added support to benchmark mixed-precision SCALV APIs(BLAS and CBLAS) - Updated the existing benchmarking file for SCALV API, to include support to call the BLAS and CBLAS mixed-precision SCALV, namely cblas_csscalv(), csscalv_(), cblas_zdscalv(), zdscalv_(). - The input is expected to be given with the datatype 'ZD' and 'CS' in order to benchmark the associated mixed-precision APIs. AMD-Internal: [CPUPL-4722] Change-Id: I4ab0fb19fe1949468cf707d0a857e8a1681addeb	2024-03-08 04:54:30 -05:00
Nallani Bhaskar	799a456abc	Fixed corner case issue in aocl_gemm addon Description 1. when mr0=1 case the accumulator register and operand registers for an fma instruction got swapped. Corrected the copy paste error. 2. Removed fill array for c_ref in bench_lpgemm.c and used memcpy from c buf, because fill array now using rand() function to initialize data which can be different when c_ref and c called separately, this was working because data was fixed (i=0 ... i%5). Change-Id: Ia513331ba49d28adc7bcdc0ec78d443abe66780b	2024-03-08 04:10:19 -05:00
Bhaskar Nallani	2ce47e6f5e	Implemented optimal AVX512-variant of f32 LPGEMV 1. The 5 LOOP LPGEMM path is in-efficient when A or B is a vector (i.e, m == 1 or n == 1). 2. An efficient implementation of lpgemv_rowvar_f32 is developed considering the b matrix reorder in case of m=1 and post-ops fusion. 3. When m = 1 the algorithm divide the GEMM workload in n dimension intelligently at a granularity of NR. Each thread work on A:1xk B:kx(>=NR) and produce C=1x(>NR). K is unrolled by 4 along with remainder loop. 4. When n = 1 the algorithm divide the GEMM workload in m dimension intelligently at a granularity of MR. Each thread work on A:(>=MR)xk B:kx1 and produce C = (>=MR)x1. When n=1 reordering of B is avoided to efficiently process in n one kernel. 5. Fixed few warnings while loading 2 f32 bias elements using _mm_load_sd using float pointer. Typecasted to (const double *) AMD-Internal: [SWLCSG-2391, SWLCSG-2353] Change-Id: If1d0b8d59e0278f5f16b499de1d629e63da5b599	2024-03-04 23:53:23 +05:30
mkadavil	d00e84ced3	Matrix Add post-operation support for float(bf16\|f32) LPGEMM APIs. -This post-operation computes C = (betaC + alphaA*B) + D, where D is a matrix with dimensions and data type the same as that of C matrix. AMD-Internal: [SWLCSG-2424] Change-Id: I9464d1f514e3b04275fe93441489b4503a08937a	2024-02-23 02:02:33 -05:00
mkadavil	01b7f8c945	Matrix Add post-operation support for integer(s16\|s32) LPGEMM APIs. -This post-operation computes C = (betaC + alphaA*B) + D, where D is a matrix with dimensions and data type the same as that of C matrix. -For clang compilers (including aocc), -march=znver1 is not enabled for zen kernels. Have updated CKVECFLAGS to capture the same. AMD-Internal: [SWLCSG-2424] Change-Id: Ie369f7ea5c80ab69eea3f3e03a8d9546e14f5c09	2024-02-12 23:51:36 +05:30
jagar	40b1af4c3f	CMake:Added cmake for bench CMakelists.txt is added in bench. Steps are provided to build for different targets. AMD-Internal: [CPUPL-2748] Change-Id: I58027f4e42d1323cafb151224c45868bc8337ff4	2024-02-06 06:50:34 -05:00
mkadavil	864170f5cb	Scalar value support for zero-point and scale-factor. -As it stands, in LPGEMM, users are expected to pass an array of values with length the same as N dimension as inputs for zero point or scale factor. However at times, a single scalar value is used as zero point or scale factor for the entire downscaling operation. The mandate to pass an array requires the user to allocate extra memory and fill it with the scalar value so as to be used in downscaling. This limitation is lifted as part of this commit, and now scalar values can be passed as zero point or scale factor. -LPGEMM bench enhancements along with new input format to improve readability as well as flexibility. AMD-Internal: [SWLCSG-2581] Change-Id: Ibd0d89f03e1acadd099382dffcabfec324ceb50f	2024-01-12 04:37:35 +05:30
Meghana Vankadari	6567df7b12	bf16bf16f32o<bf16\|f32> Fix for scaling issue when transA is enabled. Details: - LPGEMM uses bli_pba_acquire_m with BLIS_BUFFER_FOR_A_BLOCK to checkout memory when A matrix needs to be packed. This multi-threaded lock overhead becomes prominent when m/n dimensions are relatively small, even when k is large. In order to address this, bli_pba_acquire_m is used with BLIS_BUFFER_FOR_GEN_USE for LPGEMM. For *GEN_USE, the memory is allocated using aligned malloc instead of checking out from memory pool. Experiments have shown malloc costs to be far lower than memory pool guarded by locks, especially for higher thread count. - Deleted few unnecessary instructions from packing kernels. - Replaced bench_input.txt with lesser number of inputs. AMD-Internal: [CPUPL-4329] Change-Id: I5982a0a4df9dc72fab0cffab795c23822d5c8774	2023-12-21 04:53:32 +05:30
Edward Smyth	ed5010d65b	Code cleanup: AMD copyright notice Standardize format of AMD copyright notice. AMD-Internal: [CPUPL-3519] Change-Id: I98530e58138765e5cd5bc0c97500506801eb0bf0	2023-11-23 08:54:31 -05:00
Edward Smyth	f471615c66	Code cleanup: No newline at end of file Some text files were missing a newline at the end of the file. One has been added. AMD-Internal: [CPUPL-3519] Change-Id: I4b00876b1230b036723d6b56755c6ca844a7ffce	2023-11-22 17:11:10 -05:00
Meghana Vankadari	77bd9a7f17	Added parameter checking for LPGEMM APIs Change-Id: I6ea89fd0d2516539e5a4e9cd8537570b23194d89	2023-11-09 21:50:55 -05:00
Meghana Vankadari	0c12b72651	LPGEMM bench enhancements Details: - Moved the downscale & postop options from commmandline to input file. - Now the format of the input file is as follows: dt_in dt_out stor transa transb op_a op_b m n k lda ldb ldc postops - In case of no-postops, 'none' has to be passed in the place of postops. - Removed duplication of mat_mul_bench_main function for bf16 APIs. - Added a function called print_matrix for each datatype which can help in printing matrices while debugging. - Added printing of ref, computed and diff values while reporting failure. - Added new functions for memory allocation and freeing. Different types of memory allocation is chosen based on mode bench is running(performance or accuracy mode). Change-Id: Ia7d740c53035bc76e578a03869590c9f04396b72	2023-11-09 03:55:10 -05:00
Eashan Dash	c3d1a3878c	Parallelized Pack and Compute Extension APIs 1. OpenMP based multi-threading parallelism is added for BLAS extension APIs of Pack and Compute 2. Both pack and compute APIs are parallelized. 3. Multi-threading of pack and compute APIs done with different number of threads can lead to inconsistent results due to output difference of the full packed matrix buffer when packed with different number of threads. 4. In multi-threaded execution, we ensure output of packed buffer is exactly the same as in single threaded execution. 5. Similarly for compute API, read of packed buffer in multi- threaded execution is exactly the same as in single-threaded execution. 6. Routines are added to compute the offsets for thread workload distribution for MT execution. 1. The offsets are calculated in such a way that it resembles the reorder buffer traversal in single threaded reordering. 2. The panel boundaries (KCxNC) remain as it is accessed in single thread, and as a consequence a thread with jc_start inside the panel cannot consider NC range for reorder. 3. It has to work with NC' < NC, and the offset is calulated using prev NC panels spanning k dim + cur NC panel spaning pc loop cur iteration + (NC - NC') spanning current kc0 (<= KC). 7. Routines to ensure the same are added for MT execution 1. frame/base/bli_pack_compute_utils.c 2. frame/base/bli_pack_compute_utils.h AMD-Internal: [CPUPL-3560] Change-Id: I0dad33e0062519de807c32f6071e61fba976d9ac	2023-11-03 08:47:17 -04:00
Meghana Vankadari	f8f4343b55	Updated cntx with packA function pointer for AVX512_VNNI support Details: - Modified bench to support testing for sizes where matrix strides are larger than the corresponding dimensions. - Modified early-return checks in all interface APIs to check validity of strides in relation to the corresponding dimension rather than checking if strides are equal to dimensions. Change-Id: I382529b636a4acc75f6d93d997af22a168a7bfc4	2023-11-03 04:50:00 -04:00
Nallani Bhaskar	b3391ef5da	Updated ERF threshold and packa changes in bf16 Description: 1. Updated ERF function threshold from 3.91920590400 to 3.553 to match with the reference erf float implementation which reduced errors a the borders and also clipped the output to 1.0 2. Updated packa function call with pack function ptr in bf16 api to avoid compilation issues for non avx512bf16 archs 3. Updated lpgemm bench [AMD-Internal: SWLCSG-2423 ] Change-Id: Id432c0669521285e6e6a151739d9a72a7340381d	2023-10-29 23:55:46 +05:30
Meghana Vankadari	ac3e8ff01b	Bug fix and enhancements in bf16bf16f32obf16\|f32 Details: - Updated pack function call in ic loop to accept correct params. - Modified documentation in bench file to reflect updated usage of bench for downscaled APIs. - Modified memory allocation for C panel in BF16 APIs to use BLIS_BUFFER_FOR_GEN_USE while requesting for memory from pool. Change-Id: Id624ed92ae7c8dafd7f6a32fc1554d2357de4df5	2023-10-25 23:28:31 +05:30
mkadavil	26d1ab5ebc	<u\|s>8s8s<16\|32>os8 memory allocation fix to circumvent scaling issue. -When bli_pba_acquire_m api is used for packbuf type BLIS_BUFFER_FOR_ <A_BLOCK\|B_PANEL\|C_PANEL>, the memory is allocated by checking out a block from an internal memory pool. In order to ensure thread safety, the memory pool checkout is protected using mutex (bli_pba_lock/ bli_pba_unlock). When the number of threads trying to checkout memory (in parallel) are high, these locks tend to become a scaling bottleneck, especially when the memory is to be used for non-packing purposes (packing could hide some of this cost). LPGEMM uses bli_pba_acquire_m with BLIS_BUFFER_FOR_C_PANEL to checkout memory when downscale is enabled for temporary C accumulation. This multi-threaded lock overhead becomes prominent when m/n dimensions are relatively small, even when k is large. In order to address this, bli_pba_acquire_m is used with BLIS_BUFFER_FOR_GEN_USE for LPGEMM. For *GEN_USE, the memory is allocated using aligned malloc instead of checking out from memory pool. Experiments have shown malloc costs to be far lower than memory pool guarded by locks, especially for higher thread count. -LPGEMM bench fixes for crash observed when benchmarking with post-ops enabled and no downscale. AMD-Internal: [SWLCSG-2354] Change-Id: I4e92feadd2cf638bb26dd03b773556800a1a3d50	2023-10-23 10:00:32 -04:00
Arnav Sharma	c8f14edcf5	BLAS Extension API - ?gemm_compute() - Added support for 2 new APIs: 1. sgemm_compute() 2. dgemm_compute() These are dependent on the ?gemm_pack_get_size() and ?gemm_pack() APIs. - ?gemm_compute() takes the packed matrix buffer (represented by the packed matrix identifier) and performs the GEMM operation: C := A * B + beta * C. - Whenever the kernel storage preference and the matrix storage scheme isn't matching, and the respective matrix being loaded isn't packed either, on-the-go packing has been enabled for such cases to pack that matrix. - Note: If both the matrices are packed using the ?gemm_pack() API, it is the responsibility of the user to pack only one matrix with alpha scalar and the other with a unit scalar. - Note: Support is presently limited to Single Thread only. Both, pack and compute APIs are forced to take n_threads=1. AMD-Internal: [CPUPL-3560] Change-Id: I825d98a0a5038d31668d2a4b84b3ccc204e6c158	2023-10-16 08:18:52 -04:00
Meghana Vankadari	eb5ab3f762	LPGEMM: Added transB support for bf16bf16f32o<bf16\|f32> APIs Details: - Modified aocl_get_reorder_buf_size_ and aocl_reorder_ APIs to allow reordering from column major input matrix. - Added new pack kernels that packs/reorders B matrix from column-major input format. - Updated Early-return check conditions to account for trans parameters. - Updated bench file to test/benchmark transpose support. AMD-Internal: [CPUPL-2268] Change-Id: Ida66d7e3033c52cca0229c6b78d16976fbbecc4c	2023-10-12 23:36:18 +05:30
mkadavil	ea0324ab95	Multi data type downscaling support for u8s8s16 - u8s8s16<u8\|s8> Downscaling is used when GEMM output is accumulated at a higher precision and needs to be converted to a lower precision afterwards. Currently the u8s8s16 flavor of api only supports downscaling to s8 (int8_t) via aocl_gemm_u8s8s16os8 after results are accumulated at int16_t. LPGEMM is modified to support downscaling to different data types, like u8, s16, apart from s8. The framework (5 loop) passes the downscale data type to the micro-kernels. Within the micro-kernel, based on the downscale type, appropriate beta scaling and output buffer store logic is executed. This support is only enabled for u8s8s16 flavor of api's. The LPGEMM bench is also modified to support passing downscale data type for performance and accuracy testing. AMD-Internal: [SWLCSG-2313] Change-Id: I723d0802baf8649e5e41236b239880a6043bfd30	2023-10-12 09:19:56 -04:00
Meghana Vankadari	4874895a68	LPGEMM: Added transA support for bf16bf16f32o<bf16\|f32> APIs Details: - Added new params(order, trans) to aocl_get_reorder_buf_size_ and aocl_reorder_ APIs. - Added new pack kernels that packs A matrix from either row-major or column major input matrix to pack buffer with row-major format. - Updated cntx with pack kernel function pointers for packing A matrix. - Transpose of A matrix is handled by packing A matrix to row-major format during run-time. - Updated Early-return check conditions to account for trans parameters. - Updated bench file to test/benchmark transpose support. AMD-Internal: [SWLCSG-2268, SWLCSG-2442] Change-Id: I43a113dc4bc11e6bb7cc4d768c239a16cb6bbea4	2023-10-11 07:16:08 -04:00
mkadavil	c3b97559c1	Zero Point support for <u\|s>8s8s<32\|16>os8 LPGEMM APIs -Downscaled / quantized value is calculated using the formula x' = (x / scale_factor) + zero_point. As it stands, the micro-kernels for these APIs only support scaling. Zero point addition is implemented as part of this commit, with it being fused as part of the downscale post-op in the micro-kernel. The zero point input is a vector of int8 values, and currently only vector based zero point addition is supported. -Bench enhancements to test/benchmark zero point addition. AMD-Internal: [SWLCSG-2332] Change-Id: I96b4b1e5a384a4683b50ca310dcfb63debb1ebea	2023-10-10 12:05:47 +05:30
Kiran Varaganti	db4fbfe9a6	Fix compiler error for "inline" functions in LPGEMM bench Application Functions which are declared as "inline" may trigger compiler error "undefined function" This linker error is eliminated by use "static" before "inline". Therefore added "static" before all inline functions. Change-Id: I5952fb71112fc4792011c3e29be930ccfbce4562	2023-09-27 02:26:23 -04:00
mkadavil	e5e9127a68	Fixes for aocl_gemm addon compilation issues Certain functions were updated recently and now takes extra arguments for error handling. Usage of the same are now updated in aocl_gemm. Change-Id: I7daca4fd1f284d57034d564f0a08cc6410ccfd5c	2023-09-06 16:00:34 +05:30
Eleni Vlachopoulou	660cd6d1b2	Adding nrm2 target for benchmarking on Windows. Modifying blis/bench/CMakeLists.txt to include nrm2 target and produce the corresponding executable. AMD-Internal: [CPUPL-3625] Change-Id: I7945416142e07ac99510ed9500a2c620053c7e13	2023-07-10 14:03:05 -04:00
mkadavil	b167e47091	LPGEMM frame and micro-kernel updates to fix gcc9.4 compilation issue. -Micro-kernel: Some AVX512 intrinsics(eg: _mm512_loadu_epi32) were introduced in later versions of gcc (>10) in addition to already existing masked intrinsic(eg: _mm512_mask_loadu_epi32). In order to support compilation using gcc 9.4, either the masked intrinsic or other gcc 9.4 compatible intrinsic needs to be used (eg: _mm512_loadu_si512) in LPGEMM Zen4 micro-kernels. -Frame: BF16 LPGEMM api's (aocl_gemm_bf16bf16f32obf16/bf16bf16f32of32) needs to be disabled if aocl_gemm (LPGEMM) addon is compiled using gcc 9.4. BF16 intrinsics are not supported in gcc 9.4, and the micro-kernels for BF16 LPGEMM is excluded from compilation based on GNUC macro. AMD-Internal: [CPUPL-3396] Change-Id: I096b05cdceea77e3e7fec18a5e41feccdf47f0e7	2023-05-11 18:00:18 +05:30
Edward Smyth	7e50ba669b	Code cleanup: No newline at end of file Some text files were missing a newline at the end of the file. One has been added. Also correct file format of windows/tests/inputs.yaml, which was missed in commit `0f0277e104` AMD-Internal: [CPUPL-2870] Change-Id: Icb83a4a27033dc0ff325cb84a1cf399e953ec549	2023-04-21 10:02:48 -04:00
eashdash	a72fff2be9	Added NEW LPGEMM TYPE- s8s8s16os16 and s8s8s16os8 1. New LPGEMM type - s8s8s16os16 and s8s8s16os8 are added. 2. New interface, frame and kernel files are added. 3. Frame and kernel level files added and modified for s8s8s16 4. s8s8s16 type involves design changes of 2 operations - Pack B and Mat Mul 5. Pack B kernel routines to pack B matrix for s16 FMA and compute the sum of every column of B matrix to implement the s8s8s16 operation using the s16 FMA instructions. 5. Mat Mul Kernel files to compute the GEMM output using s16 FMA. Here the A matrix elements are converted from int8 to uint8 (s16 FMA works with A matrix type uint8 only) by adding extra 128 to every A matrix element 6. Post GEMM computation, additional operations are performed on the accumulated outputs to get the correct results. Final C = C - ( (sum of column of B matrix) * 128 ) This is done to compensate for the addition of extra 128 to every A matrix elements 7. With this change, two new LPGEMM APIs are introduced in LPGEMM - s8s8s16os16 and s8s8s16os8. 8. All previously added post-ops are supported on s8s8os16/os8 also. AMD-Internal: [CPUPL-3234] Change-Id: I3cc23e3dcf27f215151dda7c8db29b3a7505f05c	2023-04-21 05:30:38 -04:00
mkadavil	3572baa9d3	aocl_softmax_f32 api's for softmax computation as part of lpgemm. -Softmax is often used as the last activation function in a neural network - softmax(xi) = exp(xi)/(exp(x0) + exp(x1) + ... + exp(xn))). This step happens after the final low precision gemm computation, and it helps to have the softmax functionality that can be invoked as part of the lpgemm workflow. In order to support this, a new api, aocl_softmax_f32 is introduced as part of aocl_gemm. This api computes element-wise softmax of a matrix/vector of floats. This api invokes ISA specific vectorized micro-kernels (vectorized only when incx=1), and a cntx based mechanism (similar to lpgemm_cntx) is used to dispatch to the appropriate kernel. AMD-Internal: [CPUPL-3247] Change-Id: If15880360947435985fa87b6436e475571e4684a	2023-04-21 05:26:08 -04:00
mkadavil	ffa72f09cc	Support for multiple eltwise post-ops in low precision gemm. -Currently only one eltwise post-op (one of relu/prelu/gelu_tanh/ gelu_erf) is supported in the post-op struct along with bias or downscale. This setup was sufficient when only activation functions were supported as eltwise post-ops. But with the introduction of clip post-op(a type of non-activation eltwise operation), it has become necessary to extend the post-ops framework to support multiple eltwise operations, with the multiple eltwise often used in the form activation eltwise op + non-activation eltwise ops. The aocl post-op struct is modified and the post-op parser is updated to support this use case. -The lpgemm_bench is updated to support testing/benchmarking of the multiple eltwise operations use case. The function for accuracy checking is modified to support correctness testing irrespective of the order and count of post-ops. Additionally the help message is updated so as to better describe the capabilities of lpgemm_bench. AMD-Internal: [CPUPL-3244] Change-Id: If4ce8d7261d32073da8fa4757ed4f2ea0e94249f	2023-04-20 07:24:32 -04:00
mkadavil	99d10c3f88	Low precision gemm u8s8s16 downscale optimization. -Similar to downscale optimizations made for u8s8s32 gemm, the following optimizations are made to improve the downscale performance for u8s8s16 gemm: a. The store to temporary s16 buffer can be avoided when k < KC since intermediate accumulation will not required for the pc loop (only 1 iteration). The downscaled values (s8) are written directly to the output C matrix. b. Within the micro-kernel when beta != 0, the s8 data from the original C output matrix is loaded to a register, converted to s16 and beta scaling applied on it. The previous design of copying the s8 value to the s16 temporary buffer inside jc loop and using the same in beta scaling is removed. -Alpha scaling (multiply instruction) by default was resulting in performance regression when k dimension is small and alpha=1 in s16 micro-kernels. Alpha scaling is now only done when alpha != 1. AMD-Internal: [CPUPL-3237] Change-Id: If25f9d1de8b9b8ffbe1bd7bce3b7b0b5094e51ef	2023-04-19 06:40:06 -04:00
mkadavil	e23765010d	aocl_gelu_<tanh\|erf>_f32 api's for gelu computation as part of lpgemm. -Currently in aocl_gemm, gelu (both tanh and erf based) computation is only supported as a post-op as part of low precision gemm api call (done at micro-kernel level). However gelu computation alone without gemm is required in certain cases for users of aocl_gemm. -In order to support this, two new api's - aocl_gelu_tanh_f32 and aocl_gelu_erf_f32 are introduced as part of aocl_gemm. These api's computes element-wise gelu_tanh and gelu_erf respectively of a matrix/ vector of floats. Both the api's invokes ISA specific vectorized micro- kernels (vectorized only when incx=1), and a cntx based mechanism (similar to lpgemm_cntx) is used to dispatch to the appropriate kernel. AMD-Internal: [CPUPL-3218] Change-Id: Ifebbaf5566d7462288a9a67f479104268b0cc704	2023-04-17 05:15:56 -04:00
eashdash	12c97021a1	Added New Post-Op - Custom Clipping for LPGEMM and SGEMM 1. Custom Clip is an element-wise post-op which is used to clip the accumulated GEMM output within a certain range. 2. The Clip Post-Op is used in downscaled and non-downscaled LPGEMM APIs and SGEMM. 3. Changes are done at frame and microkernel level to implement this post-op. 4. Different versions are implemented - AVX-512, AVX-2, SSE-2 to enable custom clipping for various LPGEMM types and SGEMM AMD-Internal: [CPUPL-3207] Change-Id: I71c60be69e5a0dc47ca9336d58181c097b9aa0c6	2023-04-17 04:38:20 -04:00

1 2 3

105 Commits