- Adjusted the macro-guards for variables specific to
multithreading, when BLIS is configured with OpenMP.
- This included calling the single-threaded kernel directly
if increment is 0 as well, since this would remove an
unnecessary dependency on one of the variables used only
when we enable OpenMP.
- Further updated the condition to pack the vector, to
avoid it when increment is 0. In this case, we directly
call the kernel.
AMD-Internal: [CPUPL-5480]
Change-Id: I31a9c6e3ffc3c4f9d5b03ed8745919ad65c99c79