amd/blis - blis - Public git mirror

amd/blis

mirror of https://github.com/amd/blis.git synced 2026-04-20 07:38:53 +00:00

Author	SHA1	Message	Date
Mithun Mohan	4a95f44d39	Buffer scale support for matrix add and matrix mul post-ops in bf16 API. -Currently the values (from buffer) are added/multiplied as is to the output registers while performing the matrix add/mul post-ops. Support is added for scaling these values before using them in the post-ops. Both scalar and vector scale_factors are supported. AMD-Internal: [SWLCSG-3181] Change-Id: Ifdb7160a1ea4f5ecccfa3ef31ecfed432898c14d	2025-01-08 10:35:50 +00:00
Mithun Mohan	8d8a8e2f19	Light-weight logging framewok for LPGEMM. -A light-weight mechanism/framework to log input details and a stringified version of the post-ops structure is added to LPGEMM. Additionally the runtime of the API is also logged. The logging framework logs to a file with filename following the format aocl_gemm_log_<PID>_<TID>.txt. -To enable this feature, the AOCL_LPGEMM_LOGGER_SUPPORT=1 macro needs to be defined when compiling BLIS (with aocl_gemm addon enabled) by passing CFLAGS="-DAOCL_LPGEMM_LOGGER_SUPPORT=1" to ./configure. Additionally AOCL_ENABLE_LPGEMM_LOGGER=1 has to be exported in the environment during LPGEMM runtime. AMD-Internal: [SWLCSG-3280] Change-Id: I30bfb35b2dc412df70044601b335938fc9f49cfb	2025-01-03 11:28:57 +00:00
Meghana Vankadari	bfc512d3e1	Implemented batch_gemm for bf16bf16f32of32\|bf16 Details: - The batch matmul performs a series of matmuls, processing more than one GEMM problem at once. - Introduced a new parameter called batch_size for the user to indicate number of GEMM problems in a batch/group. - This operation supports processing GEMM problems with different parameters including dims,post-ops,stor-schemes etc., - This operation is optimized for problems where all the GEMMs in a batch are of same size and shape. - For now, the threads are distributed among different GEMM problems equally irrespective of their dimensions which leads to better performance for batches with identical GEMMs but performs sub-optimally for batches with non-identical GEMMs. - Optimizations for batches with non-identical GEMMs is in progress. - Added bench and input files for batch_matmul. AMD-Internal: [SWLCSG-2944] Change-Id: Idc59db5b8c5794bf19f6f86bcb8455cd2599c155	2025-01-03 03:28:32 -05:00
Deepak Negi	615789e196	Fixed compilation issue with clang 18 on windows Description -In enum AOCL_PARAMS_STORAGE_TYPES the member FLOAT was declared and the clang 18 compiler in msvc throwing issue with multiple definition. We replace FLOAT and BFLOAT16 to AOCL_GEMM_<F32/BF16>. AMD-Internal: CPUPL-6174 Change-Id: Ic061af068854d51629b82b495efd0eb54543f329	2024-12-12 06:37:06 -05:00
Deepak Negi	baeebe75c9	Support for standard AutoAWQ storage format. Description: 1. AutoAWQ use a int32 buffer to store 8 elements each of 4 bits in this format [0, 2, 4, 6, 1, 3, 5, 7]. 2. Support is added to convert above format back to the original sequential order [0, 1, 2, 3, 4, 5, 6, 7] before reordering in the AWQ API. AMD-Internal: SWLCSG-3169 Change-Id: I5395766060c200ab81d0b8be94356678a169ac13	2024-12-02 04:02:27 -05:00
Meghana Vankadari	fbb72d047f	Added group quantization and zero-point support for WOQ kernels Description: 1. Added group quantization and zero-point (zp) in aocl_gemm_bf16s4f32o<bf16\|f32> API. 2. Group quantization is technique to improve accuracy where scale factors to dequantize weights varies at group level instead of per channel and per tensor level. 3. Added zp and scaling in woq packb kernels so that for large M values zp and scaling are performed at pack-b stage and bf16 kernels are called 4. Adding zp support and scaling to default path in WoQ kernels created some performance overhead when M value is very small. 5. Added string group_size to lpgemm bench to read group size from bench_input.txt and tested for various combinations of matrix dimensions. 6. The scalefactors could be of type float or bf16 and the zeropoint values are expected to be in int8 format. AMD-Internal: [SWLCSG-3168, SWLCSG-3172] Change-Id: Iff07b54d76edc7408eb2ea0b29ce8b4a04a38f57	2024-12-02 06:46:13 +00:00
Deepak Negi	04ae01aeab	Added support to specify bias data type in bf16 API's Description: 1. The bias type was supported only based on output data type. 2. The option is added in the pre-ops structure to select the bias data type irrespective of the storage data type in bf16 and WoQ API's AMD-Internal: SWLCSG-3171 Change-Id: Iac10b946c2d4a5c405b2dc857362be0058615abf	2024-11-19 05:30:02 -05:00
Deepak Negi	60a8c71a1a	Sigmoid and Tanh post-operation support for int8 API's. Description: Implemented sigmoid, tanh as fused post-ops in aocl_gemm_<s8\|u8>s8<s32\|s16>o<s8\|u8\|s32> API's Sigmoid(x) = 1/1+e^(-x) Tanh(x) = (1-e^(-2x))/(1+e^(2x)) Updated bench_lpgemm to recognize sigmod, tanh as options for post-ops from bench_input and verified. AMD-Internal: [SWLCSG-3178] Change-Id: I9df3aab02222f728ff9d1f292c7bc549f30176f0	2024-11-15 05:36:31 -05:00
Deepak Negi	146f3b2eb2	Sigmoid and Tanh post-operation support for f32 API. Description: Implemented sigmoid, tanh as fused post-ops in aocl_gemm_f32f32f32of32 API's Sigmoid(x) = 1/1+e^(-x) Tanh(x) = (1-e^(-2x))/(1+e^(2x)) Updated bench_lpgemm to recognize sigmod, tanh as options for post-ops from bench_input and verified. AMD-Internal: [SWLCSG-3178] Change-Id: Iac0a907f6dea1d9cb82d9fd8716bfdbf1c33921d	2024-11-15 04:20:20 -04:00
Deepak Negi	b5c1b6055a	Sigmoid and Tanh post-operation support for bf16 API. Description: Implemented sigmoid, tanh as fused post-ops in aocl_gemm_bf16bf16f32o<f32\|bf16) API's Sigmoid(x) = 1/1+e^(-x) Tanh(x) = (1-e^(-2x))/(1+e^(2x)) Updated bench_lpgemm to recognize sigmod, tanh as options for post-ops from bench_input and verified. AMD-Internal: [SWLCSG-3178] Change-Id: I78a3ba4a67ab63f9d671fbe315f977b016a0d969	2024-11-15 01:13:31 -04:00
Mithun Mohan	097cda9f9e	Adding support for AOCL_ENABLE_INSTRUCTIONS for f32 LPGEMM API. -Currently lpgemm sets the context (block sizes and micro-kernels) based on the ISA of the machine it is being executed on. However this approach does not give the flexibility to select a different context at runtime. In order to enable runtime selection of context, the context initialization is modified to read the AOCL_ENABLE_INSTRUCTIONS env variable and set the context based on the same. As part of this commit, only f32 context selection is enabled. -Bug fixes in scale ops in f32 micro-kernels and GEMV path selection. -Added vectorized f32 packing kernels for NR=16(AVX2) and NR=64(AVX512). This is only for B matrix and helps remove dependency of f32 lpgemm api on the BLIS packing framework. AMD Internal: [CPUPL-5959] Change-Id: I4b459aaf33c54423952f89905ba43cf119ce20f6	2024-10-30 08:52:22 +00:00
Meghana Vankadari	b04b8f22c9	Introduced un-reorder API for bf16bf16f32of32 Details: - Added a new API called unreorder that converts a matrix from reordered format to it's original format( row-major or col-major ). - Currently this API only supports bf16 datatype. - Added corresponding bench and input file to test accuracy of the API. - The new API is only supported for 'B' matrix. - Modified input validation checks in reorder API to account for row Vs col storage of matrix and transposes for bf16 datatype. Change-Id: Ifb9c53b7e6da6f607939c164eb016e82514581b7	2024-10-23 07:49:24 -04:00
varshav2	dabfdf484a	Add Scale post-op for F32 API - Implemented the Scale post-op for the F32 API for all kernels - f32_scale = (f32 * scale_factor) + offset - Added the bench inputs Change-Id: Ib0f25f870eafe695d8b2a2c434c8cb3ec4f7db4c	2024-10-21 06:08:31 -04:00
Deepak Negi	16653ed208	Added support for column major B matrix in BF16S4F32F32 reorder API. -Added new pack kernels that packs/reorders B matrix from column-major input format. This also supports the transB scenario if input B matrix is row major. Change-Id: I4c75b6e81016331fd7e7f95ad4212e6d38dc586f	2024-09-20 01:11:21 +05:30
Chandrashekara K R	e4eed817aa	Added logic to use right format specifier to read integer value. Updated logic to use "%ld" and "%lld" format specifiers to read 64-bit integer from input files using fscanf function on Linux and Windows respectively when the user set INT_SIZE='auto' on 64-bit machine or INT_SIZE='64'. Otherwise "%d" on both windows and Linux for benchmarking blis and LPGEMM. Change-Id: I4762c4c1b3fcd09cf66d0cc9572d38766be6be60	2024-09-17 04:48:59 -04:00
Chandrashekara K R	91d4337b8b	Updated format specifier for fscanf to read double values. Updated format specifier to read signed double("%lld") and unsigned double("%llu") from file using fscanf from both windows and Linux. AMD-Internal: [CPUPL-5787] Change-Id: Ibef50b0df708f474e22f703240e264eff1de3994	2024-09-13 14:57:28 +05:30
Mithun Mohan	453c9f0084	Fixes for bfloat16 accumulation rounding errors in bench. For the bf16bf16of32bf16 lpgemm api, inside the micro-kernels in order to convert the accumulated float values to bfloat16 before storing, the _mm512_cvtneps_pbh intrinsic (vcvtneps2bf16) is used. This intrinsic rounds the value based on a rounding bias logic. Replicating the same rounding logic inside the bf16 bench accuracy check function to get proper one to one comparison of output values. AMD Internal: [SWLCSG-2948] Change-Id: I135ac39ac8484769b6c0fe5b3e351dd22d7ca1d8	2024-09-11 01:39:11 -04:00
Meghana Vankadari	687abe4c96	Bug fix in WOQ kernel for m=4 case. - Updated pre_op_off computation for nr0 < NR cases. - Fixed warnings in bench file. Change-Id: Iae30fa84b6b47ebd94ab05d2139056aee24546d7	2024-09-05 05:00:30 +00:00
Meghana Vankadari	2e1cc2f14a	Added bf16s4f32 kernels to handle m=4 cases Details: - In WOQ, if m = 4, special case kernels are added where s4->bf16 conversion happens inside the compute kernel and packing is avoided. For all other cases, B matrix is dequantized and packed at KC loop level and native bf16 kernels are re-used at compute level. - Fixes in bench to avoid accuracy failures when datatype of output is bf16. Change-Id: Ie8db42da536891693d5e82a5336b66514a50ccb2	2024-09-04 07:36:57 -04:00
Deepak Negi	e429e57b53	Replaced int_32 with dim_t in lpgemm bench Replaced int32_t with dim_t (int64_t) to avoid overflow. Change-Id: I4132b72fcbffd9dbd2242b3638922931bcdb1b80	2024-09-02 09:03:02 -04:00
Deepak Negi	6dcf500703	Element wise operations API for float(f32) input matrix in LPGEMM. This API supports applying element wise operations (eg: post-ops) on a float(f32) input matrix to get an output matrix of the same (float(f32)). Change-Id: I387a544f0d33d2231f5f6a92e212f17b1103dd24 AMD Internal: [SWLCSG-2947] Change-Id: I387a544f0d33d2231f5f6a92e212f17b1103dd24	2024-08-27 03:28:52 -04:00
varshav2	e3c434080a	Fix duplicate check and early return in s8s8s32/u8s8s32 - removed the duplicate check for col-major inputs in s8s8s32/u8s8s32 APIs - Fixed the print in bench_lpgemm Change-Id: If40837b89927dd82d8aa6f620d1a7f2c24aed53c	2024-08-23 02:32:20 +05:30
Edward Smyth	82bdf7c8c7	Code cleanup: Copyright notices - Standardize formatting (spacing etc). - Add full copyright to cmake files (excluding .json) - Correct copyright and disclaimer text for frame and zen, skx and a couple of other kernels to cover all contributors, as is commonly used in other files. - Fixed some typos and missing lines in copyright statements. AMD-Internal: [CPUPL-4415] Change-Id: Ib248bb6033c4d0b408773cf0e2a2cda6c2a74371	2024-08-05 15:35:08 -04:00
Edward Smyth	591a3a7395	Code cleanup: file formats and permissions - Remove execute file permission from source and make files. - dos2unix conversion. - Add missing eol at end of files. Also update .gitignore to not exclude build directory but to exclude any build_* created by cmake builds. AMD-Internal: [CPUPL-4415] Change-Id: I5403290d49fe212659a8015d5e94281fe41eb124	2024-08-05 11:52:33 -04:00
mkadavil	9f5fec7713	Matrix MUL op support in element wise operations API for bfloat16. -Matrix MUL op support added in main as well as fringe bfloat16 element wise operations kernels. -Benchmarking/testing framework for the same is added. -Fixed issues in setting up post-ops node index. AMD Internal: [SWLCSG-2947, SWLCSG-2953] Change-Id: Iba7561a6a60df41211efbf06fab1b4900207bcf8	2024-08-05 08:29:42 +05:30
Deepak Negi	80bf6249f0	Matrix MUL post-operation support for float(bf16\|f32) LPGEMM APIs. This post-operation computes C = (betaC + alphaAB) D, where D is a matrix with dimensions and data type the same as that of C matrix. AMD-Internal: [SWLCSG-2953] Change-Id: Id4df2ca76a8f696cb16edbd02c25f621f9a828fd	2024-08-05 08:25:32 -04:00
mkadavil	f040ba617f	Element wise operations API for bfloat16 input matrix in LPGEMM. -This API supports applying element wise operations (eg: post-ops) on a bfloat16 input matrix to get an output matrix of the same(bfloat16) or upscaled data type (float). -Benchmarking/testing framework for the same is added. AMD Internal: SWLCSG-2947 Change-Id: I43f1c269be1a1997d4912d8a3a97be5e5f3442d2	2024-08-05 07:17:08 -04:00
Meghana Vankadari	d5b4d3aa5e	Fixing control flow in aocl_gemm_bf16s4f32of32\|bf16 - Fixed framework of bf16s4f32of32 API to correct pointer updations. - Modified pre_op structure to exclude pre-op-offset. Now offset is passed as a separate parameter to the scale-pack functions. - Fixed work-distribution among threads in MT scenario. - Added Blocksizes and kernel-pointers and verified functionality for the new API. AMD-Internal: [SWLCSG-2943] Change-Id: I58fece240d62c798c880a2b2b7fa64e560cc753d	2024-07-29 05:12:09 -04:00
mkadavil	ec8c39541e	Test/benchmark framework updates to test WOQ workflow. -To enable Weight-only-Quantization (WOQ) workflow, new LPGEMM APIs have been developed where data types are A:bf16, B:int4 and C:f32/bf16. The testing and benchmarking framework for the same are added. AMD-Internal: [SWLCSG-2943] Change-Id: Icdc1d60819a23dd9f41382499d1a3c055c5edc17	2024-07-25 06:44:37 +05:30
mkadavil	d37c91dffa	Quantization (scale + zero point) support for BF16 LPGEMM api. -Quantization of f32 to bf16 (bf16 = (f32 * scale_factor) + zero_point) instead of just type conversion in aocl_gemm_bf16bf16f32obf16. -Support for multiple scale/sum/matrix_add/bias post-ops in a single LPGEMM api call. -Post-ops mask related fixes in lpgemv kernels . -Additional scale post-ops sanity checks. AMD-Internal: [SWLCSG-2945] Change-Id: I3b35cc413c176bb50bfdbd6acd4839a5ba7e94bb	2024-07-18 05:32:51 -04:00
srikanth pogula	1d7f6d414f	Bench APPs - change in Print statement for more params >Made changes in the print statements in bench files to print all the params of the individual APIs > Ex : removing tab & adding Func param "Dt\t n\t incx\t incy\t gflops\n" --> "Func Dt n incx incy gflops\n" > Ex : adding func, incx, incy params "dt_ch, n, alpha_r, alpha_i, beta_r, beta_i, gflops" --> "tmp, dt_ch, n, alpha_r, alpha_i, incx, beta_r, beta_i, incy, gflops" Change-Id: Ib5d151d7472d3f88c13a85a615a447dfa5e6b528	2024-07-11 02:04:19 -04:00
Nallani Bhaskar	d5133e4363	Fixed linking issue with bli_print_msg function Description: In recent changes bli_print_msg is used in lpgemm test application file bench_lpgemm.c for printing error message. bli_print_msg is a blis library function which is not exported for the usage of applications, because of which linking failed when blis shared library is used to build. Updated bli_print_msg with printf in the bench_lpgemm.c AMD Internal: CPUPL-5326 Change-Id: I021849baa6881bd997013e42013db1c5c711627f	2024-07-09 23:18:53 +05:30
Varaganti, Kiran	2ac24d1f9c	Avoided Extra copy of "c" matrix Initailized c_save instead of 'c" and then removed copying c to c_save. Because at the start every n_repeats iteration we are copying back c_save to c. Therefore if we initialize c_save, we can avoid extra copy of "c" to c_save before calling GEMM. For very large sizes matrix initialization takes considerable amount of time. This can be reduced now. Change-Id: I2c6ffe169e991607314897cb0c1fbfc0d74ef179	2024-07-09 00:54:03 -04:00
Meghana Vankadari	4e6fa17c08	Bug fix in LPGEMV for INT8 APIs Details: - Corrected the usage of vpdpbusd instruction in GEMV implementation for INT8 APIs. - Modified bench to fill matrices with values ranging between -5 and +5 whenever the datatype is a signed integer. Change-Id: I457462b888b667d8a34c53de762e9b4aee784ecc	2024-06-27 04:22:04 +05:30
mkadavil	a5c4a8c7e0	Int4 B matrix reordering support in LPGEMM. Support for reordering B matrix of datatype int4 as per the pack schema requirements of u8s8s32 kernel. Vectorized int4_t -> int8_t conversion implemented via leveraging the vpmultishiftqb instruction. The reordered B matrix will then be used in the u8s8s32o<s32\|s8> api. AMD-Internal: [SWLCSG-2390] Change-Id: I3a8f8aba30cac0c4828a31f1d27fa1b45ea07bba	2024-06-24 07:55:34 -04:00
Chandrashekara K R	fa75ce725e	CMake: Added logic to link openmp library given through OpenMP_libomp_LIBRARY cmake variable on linux. Enabled command line option to link libiomp5.so or libomp.so or libgomp.so libraries using cmake. Eg:- -DOpenMP_libomp_LIBRARY=<path to openmp library including library name>. If we not set above variable, by default openmp library will be libomp.so for clang and libgomp.so for gcc compiler. Change-Id: I5bffa10ff8351f5d10f0d543cbdf55aa16c84c90	2024-06-10 04:41:23 -04:00
mkadavil	cd032225ca	BF16 bias support for bf16bf16f32ob16. -As it stands the bf16bf16f32ob16 API expects bias array to be of type float. However actual use case requires the usage of bias array of bf16 type. The bf16 micro-kernels are updated to work with bf16 bias array by upscaling it to float type and then using it in the post-ops workflow. -Corrected register usage in bf16 JIT generator for bf16bf16f32ob16 API when k > KC. AMD-Internal: [SWLCSG-2604] Change-Id: I404e566ff59d1f3730b569eb8bef865cb7a3b4a1	2024-05-23 04:48:20 +05:30
Nallani Bhaskar	29db6eb42b	Added transB in all AVX512 based int8 API's Description: --Added support for tranB in u8s8s32o<s32\|s8> and s8s8s32o<s32\|s8> API's --Updated the bench_lpgemm by adding options to support transpose of B matrix --Updated data_gen_script.py in lpgemm bench according to latest input format. AMD-Internal: [SWLCSG-2582] Change-Id: I4a05cc390ae11440d6ff86da281dbafbeb907048	2024-05-23 03:46:13 +05:30
mkadavil	118e955a22	SWISH post-op support for all LPGEMM APIs. SWISH post-op computes swish(x) = x / (1 + exp(-1 * alpha * x)). SiLU = SWISH with alpha = 1. AMD-Internal: [SWLCSG-2387] Change-Id: I55f50c74a8583a515f7ea58fa0878ccbcdd6cc26	2024-05-06 06:05:11 -04:00
Vignesh Balasubramanian	1b7980a38d	Added support to benchmark AXPYV APIs - Implemented the feature to benchmark ?AXPYV APIs for the supported datatypes. The feature allows to benchmark BLAS, CBLAS or the native BLIS API, based on the macro definition. - Added a sample input file to provide examples to benchmark AXPYV for all its datatype supports. - Updated the sample input file for SCALV to provide examples to benchmark all of its datatype supports. AMD-Internal: [CPUPL-4805] Change-Id: I550920e3a57fcc2e4900e9e698330d8b8595bdee	2024-04-08 00:06:54 -04:00
Arnav Sharma	f71495a135	Support for DOTC in DOTV Bench and DTL updates - Added support for ?DOTC in bench. - Updated DTL to accept conjx as a parameter: - 'N', i.e., no conjugate for DOTU - 'C', i.e., conjugate for DOTC - Updated DTL calls in the interface with respective values of conjx. AMD-Internal: [CPUPL-4804] Change-Id: I447b19a6273566c6021c1721ce173bac4a59142c	2024-04-04 12:27:53 +05:30
jagar	bd80488af1	CMake: Update code to support blastest for ILP64 on windows Change-Id: I8e87ee073ffcb893fbcc7c9580add217ae347449	2024-03-27 12:02:26 -04:00
Eleni Vlachopoulou	020b9ff7f0	CMake: Enable builds for both static and shared builds for Linux. - Added BUILD_STATIC_LIBS option which is on by default, only on Linux. - Added TEST_WITH_SHARED option which is off by default, only on Linux. - If only shared or static lib is being built, that's the one that will be used for testing. - If both are being built, TEST_WITH_SHARED determins which library wil be used for testing. - Set linux workflows so that they build both static and shared libs, and use linux-static and linux-shared to denote which one should be used for testing. - Set -fPIC for both static and shared builds to fix issues faced when building blis using AOCC 4.0.0 and gtestsuite using gcc 9.4.0. AMD-Internal: [CPUPL-2748] Change-Id: I4227bab97ff31ecddfe218e18499f33b4e4ee63e	2024-03-14 10:32:51 -04:00
jagar	e2de45b454	CMake:Added support for ADDON(aocl_gemm) on Windows CMakelists.txt is updated to support aocl_gemm on windows. On windows, BLIS library(blis+aocl_gemm) is built successfully only with AOCC Compiler. (Clang has an issue with optimizing VNNI instructions). $cmake .. -DENABLE_ADDON="aocl_gemm" .... AMD-Internal: [CPUPL-2748] Change-Id: I9620878ab6934233fadc9ddc5d5e82ad85be9209	2024-03-14 07:57:02 -04:00
Vignesh Balasubramanian	d1a6517642	Added support to benchmark mixed-precision SCALV APIs(BLAS and CBLAS) - Updated the existing benchmarking file for SCALV API, to include support to call the BLAS and CBLAS mixed-precision SCALV, namely cblas_csscalv(), csscalv_(), cblas_zdscalv(), zdscalv_(). - The input is expected to be given with the datatype 'ZD' and 'CS' in order to benchmark the associated mixed-precision APIs. AMD-Internal: [CPUPL-4722] Change-Id: I4ab0fb19fe1949468cf707d0a857e8a1681addeb	2024-03-08 04:54:30 -05:00
Nallani Bhaskar	799a456abc	Fixed corner case issue in aocl_gemm addon Description 1. when mr0=1 case the accumulator register and operand registers for an fma instruction got swapped. Corrected the copy paste error. 2. Removed fill array for c_ref in bench_lpgemm.c and used memcpy from c buf, because fill array now using rand() function to initialize data which can be different when c_ref and c called separately, this was working because data was fixed (i=0 ... i%5). Change-Id: Ia513331ba49d28adc7bcdc0ec78d443abe66780b	2024-03-08 04:10:19 -05:00
Bhaskar Nallani	2ce47e6f5e	Implemented optimal AVX512-variant of f32 LPGEMV 1. The 5 LOOP LPGEMM path is in-efficient when A or B is a vector (i.e, m == 1 or n == 1). 2. An efficient implementation of lpgemv_rowvar_f32 is developed considering the b matrix reorder in case of m=1 and post-ops fusion. 3. When m = 1 the algorithm divide the GEMM workload in n dimension intelligently at a granularity of NR. Each thread work on A:1xk B:kx(>=NR) and produce C=1x(>NR). K is unrolled by 4 along with remainder loop. 4. When n = 1 the algorithm divide the GEMM workload in m dimension intelligently at a granularity of MR. Each thread work on A:(>=MR)xk B:kx1 and produce C = (>=MR)x1. When n=1 reordering of B is avoided to efficiently process in n one kernel. 5. Fixed few warnings while loading 2 f32 bias elements using _mm_load_sd using float pointer. Typecasted to (const double *) AMD-Internal: [SWLCSG-2391, SWLCSG-2353] Change-Id: If1d0b8d59e0278f5f16b499de1d629e63da5b599	2024-03-04 23:53:23 +05:30
mkadavil	d00e84ced3	Matrix Add post-operation support for float(bf16\|f32) LPGEMM APIs. -This post-operation computes C = (betaC + alphaA*B) + D, where D is a matrix with dimensions and data type the same as that of C matrix. AMD-Internal: [SWLCSG-2424] Change-Id: I9464d1f514e3b04275fe93441489b4503a08937a	2024-02-23 02:02:33 -05:00
mkadavil	01b7f8c945	Matrix Add post-operation support for integer(s16\|s32) LPGEMM APIs. -This post-operation computes C = (betaC + alphaA*B) + D, where D is a matrix with dimensions and data type the same as that of C matrix. -For clang compilers (including aocc), -march=znver1 is not enabled for zen kernels. Have updated CKVECFLAGS to capture the same. AMD-Internal: [SWLCSG-2424] Change-Id: Ie369f7ea5c80ab69eea3f3e03a8d9546e14f5c09	2024-02-12 23:51:36 +05:30
jagar	40b1af4c3f	CMake:Added cmake for bench CMakelists.txt is added in bench. Steps are provided to build for different targets. AMD-Internal: [CPUPL-2748] Change-Id: I58027f4e42d1323cafb151224c45868bc8337ff4	2024-02-06 06:50:34 -05:00

1 2 3 4

182 Commits