amd/blis - blis - Public git mirror

amd/blis

mirror of https://github.com/amd/blis.git synced 2026-05-20 08:29:01 +00:00

Author	SHA1	Message	Date
Meghana	c20c96d9c0	Made some critical changes to small_gemm kernels Details: - In case of GEMM, whenever beta is zero, we need to perform C = alpha (A B) instead of C = beta * C + alpha * (A * B) Added conditions to check the value of beta at different levels inside small_gemm kernels and decide whether to perform scaling C with beta or not. -Modified small_gemm kernels to use BLIS specific functions to retrieve different fields of objects. -Calling bli_gemm_check before entering bli_gemm_small to facilitate early return in case of invalid inputs. -For corner cases inside small_gemm kernels, a buffer called f_temp is used to load and store data to and from registers. populating the buffer with zeroes before use. -In bli_gemm_front, datatypes of status and return value from bli_gemm_small are not matching. Corrected the datatype of the variable 'status' inside bli_gemm_front to err_t. Change-Id: I8b52ad55008f028d6c8b7e0d20f746a869d9daea Signed-off-by: Meghana Vankadari <Meghana.Vankadari@amd.com> AMD-Internal: [CPUPL-689,SWLCSG-104]	2020-03-19 16:30:04 +05:30
Nallani Bhaskar	83745c7ffc	Beta Zero Check for sgemm small. Core Software Group SWLCSG-137 BLIS-ST validation failures Change-Id: I21d5eae6ec390438be847f2dca42350b97059d6e	2020-03-09 02:55:51 -04:00
Nallani Bhaskar	e0c95d77e1	Beta Zero Checks for sgemm_small Change-Id: I111b66ad54a27b1977d155904738a55a351e6689	2020-03-09 02:55:25 -04:00
dzambare	f965b95d8b	CPUPL-587: Corrected condition for A packing in sgemm_small Change-Id: I1e5dc4a1dbe2f1d17f9c72e8dd0c6728ac1fd750	2020-01-27 11:08:20 +05:30
Meghana	b3e2938b9e	Fix for CPUPL-549: TRSM for AlXB case results in NaN values For the kernel of size 4x8, cs_b is used instead of cs_a to calculate address of diagonal elements of matrix A. Correcting the mistake. Change-Id: Ie74e0f6a397fcd32fefb5804cd00f1e90bfe5523	2019-12-21 23:12:09 +05:30
Dipal M Zambare	72f4a7ab1e	Increased pool buffer size to accommodate packing buffers needed in small_gemm to make it reentrant. Change-Id: I96ac19ce97c39becce2c6e7ab47c3e7624560b30	2019-12-19 14:45:13 +05:30
Meghana Vankadari	62e00b4d64	Merge "Change in threshold condition for trsm_small kernels" into amd-staging-rome-rel-2.1	2019-12-17 23:54:01 -05:00
Meghana Vankadari	8eb264f78b	Change in threshold condition for trsm_small kernels Change-Id: I396e246b1639d300fcb94bdf7e5fa8bc8c87e994	2019-12-16 18:54:48 +05:30
Devrajegowda, Kiran	1fe8edbed0	"Merge Selective Packing code from amd branch flame/blis" Change-Id: Ifbdf49735f56a66fbbc96dab6d3ca6069302daed	2019-12-16 14:48:53 +05:30
Kiran Devrajegowda	21224e8264	Merge "Revert " Merge Selective Packing code from amd branch flame/blis"" into amd-staging-rome-rel-2.1	2019-12-13 00:45:34 -05:00
Nallani Bhaskar	10a26a7357	Merge "Fix for CPUPL-550: AOCC clang compiler error. Resolved: Duplicate back to back declaration of a lable in asm file" into amd-staging-rome-rel-2.1	2019-12-13 00:25:49 -05:00
Kiran Varaganti	1650bcb623	Revert " Merge Selective Packing code from amd branch flame/blis" This reverts commit `e4a6af33f5`. Reason for revert: <Review not done> Change-Id: Iae548f949a81a66281023c860c2bcffdfdae21b2	2019-12-13 00:01:35 -05:00
Nallani Bhaskar	dc4e7d1203	Fix for CPUPL-550: AOCC clang compiler error. Resolved: Duplicate back to back declaration of a lable in asm file Change-Id: I82c386d5fc00139da74fa031980d65c6a3874bd0	2019-12-12 20:43:47 +05:30
Devrajegowda, Kiran	e4a6af33f5	Merge Selective Packing code from amd branch flame/blis Change-Id: I6d577f67ec84febe6af3635b10e5c9c77844ccd2	2019-12-12 15:22:21 +05:30
Nallani Bhaskar	44edee7404	Added support to handle 7x16,8x16,9x16 efficiently in 6x16n kernel	2019-12-10 16:09:46 +05:30
Kiran Varaganti	9b6c04d075	Merge " change in threshold condition for SUP and small kernels" into amd-staging-rome-rel-2.1	2019-12-08 23:42:25 -05:00
Devrajegowda, Kiran	3192914a1c	change in threshold condition for SUP and small kernels Change-Id: I7dbd30b2004c67122a639f081efc36e0f0d69fad	2019-12-09 01:31:58 +05:30
Kiran Varaganti	27d2b5a0db	Merge "Made some improvements to trsm_small kernels" into amd-staging-rome-rel-2.1	2019-12-06 05:21:34 -05:00
Meghana	17b3a2639e	Made some improvements to trsm_small kernels Interchanged some loops to favour column-major storage. Added check condiion to identify last column and load it using a 'for' loop to avoid memory accesses out of buffer Change-Id: Id5d2e16c65017a7f4b641d33228d23903efd09ac	2019-12-06 14:48:28 +05:30
Nallani Bhaskar	af94ba29cf	Added sup support for sgemm under zen and related frame work changes. Change-Id: Ia7e88b96d3a3617e8d24754f50db081ffe2e9955	2019-12-04 10:56:10 +05:30
Meghana Vankadari	31bfe8985f	re-enabling the boundary check condition for bli_dtrsm_small_AlXB. It was disabled by mistake in previous commits. Change-Id: Ib7d2d0c5e133ff10559ce3dc5f7e624707e43c11	2019-12-03 17:07:37 +05:30
Meghana	cef185250e	Fixed Segmentation fault in trsm_small kernels for the case AlXB. For matrix sizes which are not multiples of 4, trsm_small kernels access memory outside the allocated buffers which causes segmentation fault. This is fixed by handling each of the corner cases separately. Change-Id: Ia7cfad5d65339a209a7376cc1654382593c933af	2019-12-03 17:05:57 +05:30
prangana	13249e83e2	Replace bli_thread_init_rntm with bli_rntm_init_from_global in zen small gemm Change-Id: I14fb2795b483368580ff3fcf5f537723f3845377	2019-11-30 16:33:10 +05:30
Devrajegowda, Kiran	c4047e491a	Merge branch 'amd-blis-nov-mergetest' into amd-staging-rome2.1 Change-Id: I1e04592dd9494faa34555008dd1edbca8a092a44	2019-11-29 23:01:51 +05:30
Dipal M Zambare	e6e66fb1f9	Fixed reentrancy issues with bli_sgemm_small() and bli_dgemm_small(). Replaced global buffer used for packing with the buffer provided by memory pools. These buffers are checkout at the beginning of each call and return the pool once done. Please check comment in the above functions for details. Change-Id: I76b3560f7efcc621a4455e834fce06f629c38f50	2019-11-27 19:10:16 +05:30
Meghana	c63a078a57	Fixed segemntation fault in trsm_small kernels for cases XAuB, XAltB, XAlB For matrix sizes which are not multiples of 4, trsm_small kernels access memory outside the allocated buffers which causes segmentation fault. This is fixed by handling each of the corner cases separately. Change-Id: I267e69ee095a8ca3e8ce2a3ada5f48bfefcc2219	2019-11-21 12:31:09 +05:30
Field G. Van Zee	29b0e1ef4e	Code review + tweaks to AMD's AOCL 2.0 PR (#349 ). Details: - NOTE: This is a merge commit of 'master' of git://github.com/amd/blis into 'amd-master' of flame/blis. - Fixed a bug in the downstream value of BLIS_NUM_ARCHS, which was inadvertantly not incremented when the Zen2 subconfiguration was added. - In bli_gemm_front(), added a missing conditional constraint around the call to bli_gemm_small() that ensures that the computation precision of C matches the storage precision of C. - In bli_syrk_front(), reorganized and relocated the notrans/trans logic that existed around the call to bli_syrk_small() into bli_syrk_small() to minimize the calling code footprint and also to bring that code into stylistic harmony with similar code in bli_gemm_front() and bli_trsm_front(). Also, replaced direct accessing of obj_t fields with proper accessor static functions (e.g. 'a->dim[0]' becomes 'bli_obj_length( a )'). - Added #ifdef BLIS_ENABLE_SMALL_MATRIX guard around prototypes for bli_gemm_small(), bli_syrk_small(), and bli_trsm_small(). This is strictly speaking unnecessary, but it serves as a useful visual cue to those who may be reading the files. - Removed cpp macro-protected small matrix debugging code from bli_trsm_front.c. - Added a GCC_OT_9_1_0 variable to build/config.mk.in to facilitate gcc version check for availability of -march=znver2, and added appropriate support to configure script. - Cleanups to compiler flags common to recent AMD microarchitectures in config/zen/amd_config.mk, including: removal of -march=znver1 et al. from CKVECFLAGS (since the -march flag is added within make_defs.mk); setting CRVECFLAGS similarly to CKVECFLAGS. - Cleanups to config/zen/bli_cntx_init_zen.c. - Cleanups, added comments to config/zen/make_defs.mk. - Cleanups to config/zen2/make_defs.mk, including making use of newly- added GCC_OT_9_1_0 and existing GCC_OT_6_1_0 to choose the correct set of compiler flags based on the version of gcc being used. - Reverted downstream changes to test/test_gemm.c. - Various whitespace/comment changes.	2019-10-11 10:24:24 -05:00
kdevraje	13806ba3b0	This check in has changes w.r.t Copyright information, which is changed to (start year) - 2019 Change-Id: Ide3c8f7172210b8d3538d3c36e88634ab1ba9041	2019-05-27 16:24:43 +05:30
Meghana	ee123f5358	Defined small matrix thresholds for TRSM for various cases for NAPLES and ROME Updated copyright information for kernels/zen/bli_trsm_small.c file Removed separate kernels for zen2 architecture Instead added threshold conditions in zen kernels both for ROME and NAPLES Change-Id: Ifd715731741d649b6ad16b123a86dbd6665d97e5	2019-05-27 15:36:44 +05:30
Meghana	e05171118c	Implemented TRSM for small matrices for cases where A is on the right Added separate kernels for zen and zen2 Change-Id: I6318ddc250cf82516c1aa4732718a35eae0c9134	2019-05-23 16:17:19 +05:30
kdevraje	02920f5c48	make checkblis fails for matrix dimension check at the begining hence reverting it Change-Id: Ibd2ee8c2d4914598b72003fbfc5845be9c9c1e87	2019-05-23 15:29:59 +05:30
kdevraje	84215022f2	Adding threshold condition to dgemm small matrix kernels, defining the constants in zen2 configuration Change-Id: I53a58b5d734925a6fcb8d8bea5a02ddb8971fcd5	2019-05-23 14:33:47 +05:30
kdevraje	9d76688ad9	Fix for single rank crash with HPL application. When computing offset of C buffer, as integer variables are used for a row and column index, the intermediate result value overflows and a negative value gets added to the buffer, when the negative value is too large it would index the buffer out of the range resulting in segmentation fault. Although the crash is a result of dgemm kernel, added similar code in sgemm kernel also. Change-Id: I171119b0ec0dfbd8e63f1fcd6609a94384aabd27	2019-04-11 10:23:26 +05:30
Kiran Varaganti	3a929a3d0b	Fixed code merging: bli_gemm_small.c - missed conditional checks for L!=0 && K!=0. Now they are added. This fix is done to pass blastest Change-Id: Idc9c9a04d2015a68a19553c437ecaf8f1584026c	2019-03-18 10:51:41 +05:30
Kiran Varaganti	f5ed95ecd7	Merged BLIS Release 1.3 Modified config/zen/make_defs.mk, now CKVECFLAGS := -mavx2 -mfpmath=sse -mfma -march=znver1 Change-Id: Ia0942d285a21447cd0c470de1bc021fe63e80d81	2019-03-05 15:03:57 +05:30
Field G. Van Zee	0645f239fb	Remove UT-Austin from copyright headers' clause 3. Details: - Removed explicit reference to The University of Texas at Austin in the third clause of the license comment blocks of all relevant files and replaced it with a more all-encompassing "copyright holder(s)". - Removed duplicate words ("derived") from a few kernels' license comment blocks. - Homogenized license comment block in kernels/zen/3/bli_gemm_small.c with format of all other comment blocks.	2018-12-04 14:31:06 -06:00
Field G. Van Zee	3c52725693	Renamed/moved l3 zen ukernels to haswell kernel set. Details: - Renamed the microkernels in kernels/zen/3 to kernels/haswell/3 and then updated the file contents to use the 'haswell' infix. - Updated bli_cntx_init_zen.c and bli_cntx_init_haswell.c according to above function renames. - Moved/updated the corresponding prototypes in bli_kernels_zen.h to bli_kernels_haswell.h. - Updated config_registry according to above changes. - NOTE: This rename reflects the fact that haswell microkernels are specifically written to overcome the floating-point latency for FMA instructions on Intel Haswell-like architectures, which can issue two FMA instructions per cycle. These ukernels happen to work fine on AMD Zen-based architectures. However, Zen only issues one FMA per cycle, which, while halving its floating-point throughput, gives it extra flexibility in the design of its microkernels--namely, mr and nr can be smaller and still overcome the floating-point latency for those single-issue cores. A smaller value of mr and nr allows for a larger value of kc, which may be useful in some situations. In the future, we may write such Zen-specific microkernels to take advantage of this additional flexibility.	2018-10-17 14:56:22 -05:00
Field G. Van Zee	4fa4cb0734	Trivial comment header updates. Details: - Removed four trailing spaces after "BLIS" that occurs in most files' commented-out license headers. - Added UT copyright lines to some files. (These files previously had only AMD copyright lines but were contributed to by both UT and AMD.) - In some files' copyright lines, expanded 'The University of Texas' to 'The University of Texas at Austin'. - Fixed various typos/misspellings in some license headers.	2018-08-29 18:06:41 -05:00
Field G. Van Zee	89e178ce38	Merge branch 'master' into dev	2018-07-04 17:51:16 -05:00
Isuru Fernando	14648e1376	Native windows support using clang (#227 ) * Add appveyor file * Build script * Remove fPIC for now * copy as * set CC and CXX * Change the order of immintrin.h * Fix testsuite header * Move testsuite defs to .c * Fix appveyor file * Remove fPIC again and fix strerror_r missing bug * Remove appveyor script * cd to blis directory * Fix sleep implementation * Add f2c_types_win.h * Fix f2c compilation * Remove rdp and rename appveyor.yml * Remove setenv declaration in test header * set CPICFLAGS to empty * Fix another immintrin.h issue * Escape CFLAGS and LDFLAGS * Fix more ?mmintrin.h issues * Build x86_64 in appveyor * override LIBM LIBPTHREAD AR AS * override pthreads in configure * Move windows definitions to bli_winsys.h * Fix LIBPTHREAD default value * Build intel64 in appveyor for now	2018-07-04 17:48:42 -05:00
Devin Matthews	a7166feb10	Finish macroization of assembly ukernels.	2018-06-25 12:09:18 -05:00
Devin Matthews	b4d94e54d4	Convert x86 microkernels to assembly macros.	2018-06-20 14:07:24 -05:00
Field G. Van Zee	5140ee3424	Updated types of bli_is_[un]aligned_to() functions. Details: - Changed the void* arguments of the following static functions: bli_is_aligned_to() bli_is_unaligned_to() bli_offset_past_alignment() to siz_t, and the return type of bli_offset_past_alignment() from guint_t to siz_t. This allows for more versatile usage of these functions (e.g. when aligning both pointers and leading dimension). - Updated all invocations of these functions, mostly in kernels/penryn but also in kernels/bgq, to include explicit typecasts to siz_t when pointer arguments are passed in. - Thanks to Devin Matthews for pointing out this potential bug (via issue #211). - Deleted a few trailing spaces in various penryn kernels. - Removed duplicate instances of the words "derived" and "THEORY" from various kernel license headers, likely from a malformed recursive sed performed long ago.	2018-05-23 16:56:14 -05:00
Field G. Van Zee	2e31dd7852	Inserted missing integer typecasting into ukernels. Details: - Inserted missing safeguards into most microkernels to ensure that the integers read by the microkernel's assembly instructions are of the appropriate size. In many cases, this bug was going undetected likely because the compiler was inserting zero padding before the integers in the calling function, allowing the assembly code to read 64-bits in a way that did not corrupt the "lower" 32 integer bits with garbage in the higher bits. Thanks to Francisco Igual and Devangi Parikh for finding this issue.	2018-05-16 17:28:33 -05:00
Field G. Van Zee	4b36e85be9	Converted function-like macros to static functions. Details: - Converted most C preprocessor macros in bli_param_macro_defs.h and bli_obj_macro_defs.h to static functions. - Reshuffled some functions/macros to bli_misc_macro_defs.h and also between bli_param_macro_defs.h and bli_obj_macro_defs.h. - Changed obj_t-initializing macros in bli_type_defs.h to static functions. - Removed some old references to BLIS_TWO and BLIS_MINUS_TWO from bli_constants.h. - Whitespace changes in select files (four spaces to single tab).	2018-05-08 14:26:30 -05:00
Field G. Van Zee	3defc7265c	Applied `34b72a3` to non-active/unused microkernels. Details: - Applied the read-beyond-bounds bugfix in `34b72a3` to other haswell and zen kernels (ie: other microtile shapes) which are not used by default. This was done mostly in case someone decided to pick up these kernels and start using them, not because it affects BLIS's behavior out-of-the-box.	2018-02-23 17:38:19 -06:00
Field G. Van Zee	34b72a3517	Fixed obscure read-beyond-bounds bug in sgemm ukrs. Details: - Fixed an obscure bug in the bli_sgemm_haswell_asm_6x16 and bli_sgemm_zen_asm_6x16 microkernels when the input/output matrix C is stored with general stride (ie: both rs and cs are non-unit). The bug was rooted in the way those microkernels read from matrix C-- namely, they used vmovlps/vmovhps instead of movss. By loading two floats at a time, even if one of them was treated as junk, the assembly code could be written in a more concise manner. However, under certain conditions--if m % mr == 0 and n % nr == 0 and the underlying matrix is not an internal "view" into a larger matrix-- this could result in the very last vmovhps of the last (bottom-right) microkernel invocation reading beyond valid memory. Specifically, the low 32 bits read would always be valid, but the high 32 bits could reside beyond the bounds of the array in which the output C matrix is contained. To remedy this situation, we now selectively use movss to load any element that could be the last element in the matrix.	2018-02-23 16:33:32 -06:00
Field G. Van Zee	16813335bd	Merge branch 'amd' into rt Details: - Merged contributions made by AMD via 'amd' branch (see summary below). Special thanks to AMD for their contributions to-date, especially with regard to intrinsic- and assembly-based kernels. - Added column storage output cases to microkernels in bli_gemm_zen_asm_d6x8.c and bli_gemmtrsm_l_zen_asm_d6x8.c. Even with the extra cost of transposing the microtile in registers, this is much faster than using the general storage case when the underlying matrix is column-stored. - Added s and d assembly-based zen gemmtrsm_u microkernel (including column storage optimization mentioned above). - Updated zen sub-configuration to reflect presence of new native kernels. - Temporarily reverted zen sub-configuration's level-3 cache blocksizes to smaller haswell values. - Temporarily disabled small matrix handling for zen configuration family in config/zen/bli_family_zen.h. - Updated zen CFLAGS according to changes in `1e4365b`. - Updated haswell microkernels such that: - only one vzeroupper instruction is called prior to returning - movapd/movupd are used in leiu of movaps/movups for double-real microkernels. (Note that single-real microkernels still use movaps/movups.) - Added kernel prototypes to kernels/zen/bli_kernels_zen.h, which is now included via frame/include/bli_arch_config.h. - Minor updates to bli_amaxv_ref.c (and to inlined "test" implementation in testsuite/src/test_amaxv.c). - Added early return for alpha == 0 in bli_dotxv_ref.c. - Integrated changes from `f07b176`, including a fix for undefined behavior when executing the 1m method under certain conditions. - Updated config_registry; no longer need haswell kernels for zen sub-configuration. - Tweaked marginal and pass thresholds for dotxf. - Reformatted level-1v, -1f, and -3 amd kernels and inserted additional comments. - Updated LICENSE file to explicitly mention that parts are copyright UT-Austin and AMD. - Added AMD copyright to header templates in build/templates. Summary of previous changes from 'amd' branch. - Added s and d assembly-based zen gemm microkernels (d6x8 and d8x6) and s and d assembly-based zen gemmtrsm_l microkernels (d6x8). - Added s and d intrinsics-based zen kernels for amaxv, axpyv, dotv, dotxv, and scalv, with extra-unrolling variants for axpyv and scalv. - Added a small matrix handler to bli_gemm_front(), with the handler implemented in kernels/zen/3/bli_gemm_small_matrix.c. - Added additional logic to sumsqv that first attempts to compute the sum of the squares via dotv(). If there is a floating-point exception (FE_OVERFLOW), then the previous (numerically conservative) code is used; otherwise, the result of dotv() is square-rooted and stored as the result. This new implementation is only enabled when FE_OVERFLOW is #defined. If the macro is not #defined, then the previous implementation is used. - Added axpyv and dotv standalone test drivers to test directory. - Added zen support to old cpuid_x86.c driver in build/auto-detect/old. - Added thread-local and __attribute__-related macros to bli_macro_defs.h.	2018-02-21 17:43:32 -06:00

48 Commits