mirror of
https://github.com/amd/blis.git
synced 2026-05-04 06:21:12 +00:00
Description: 1. Updated the thread partition logic for aocl_gemm_f32f32f32of32 for m<MR, n<NR cases and also balanced thread in m, n directions such that each thread gets equal amount of work and not to span thread without any work. 2. Disabled dynamic enabling of packing of a and b matrixes for smaller sizes for genoa architecture. AMD-Internal: [SWLCSG-2353 , SWLCSG-2391] Change-Id: I03b2c50e592c2e9d336ea84c0e0394af63a34cec